Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
47 changes: 47 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,52 @@
# Changelog

## 0.9.0-dev.10 — llama.cpp b10182, load-mode API, isolate log-callback fix

Native rebuild required — `src/llama.cpp` moved `d6d0ce82` → `afeebe10`
(tag `b10182`), about seven weeks of upstream work.

### Changed

- Adapted to upstream `e6dd0e29a`, which collapsed the `use_mmap` /
`use_direct_io` / `use_mlock` booleans in `llama_model_params` into a
single `llama_load_mode` enum. **`ModelParams` keeps its three
booleans** — they are mapped at the FFI boundary, so callers are
unaffected. One semantic caveat: the enum has no direct-I/O-plus-mlock
value, so when both are requested direct I/O wins and mlock is dropped.
`useDirectIo` keeps its documented precedence over `useMmap`.
- Regenerated FFI bindings against the new pin. New upstream C API now
reachable but not yet wrapped: `llama_model_n_layer_nextn`,
`llama_model_ftype`, `llama_ftype_name`,
`llama_vocab_get_suppress_tokens`, and the mtmd batch-encoding API
(`mtmd_batch_init` / `_add_chunk` / `_encode` / `_get_output_embd`).
- `mtmd_encode` is deprecated upstream in favor of `mtmd_encode_chunk`.
This package reaches multimodal via `mtmd_helper_eval_chunks` and never
called it, so no change was needed.

### Fixed

- **`LlamaLibrary.dispose` now clears the log callback.**
`LlamaLog.silence` installs a `Pointer.fromFunction` bound to the
isolate that registered it, but the slot it occupies lives in
process-global llama.cpp/ggml state and outlives that isolate. The
stale pointer stayed installed, so the next isolate to emit a log line
invoked a callback owned by a dead isolate and the VM aborted with
"Cannot invoke native callback from a different isolate". Surfaced by
Dart 3.12's stricter cross-isolate check.

Known remaining issue: parallel `dart test` still hits the concurrent
variant of this race, where one isolate holds a live callback while
another loads a model. Run the model-backed suite with `-j 1` until
`silence()` stops using a Dart callback altogether.

### Tooling

- `ffigen` 20.1.1 → 21.0.0, `lints` 5.0.0 → 6.1.0 (dev dependencies).
Note ffigen 21 requires Dart SDK ≥ 3.10 to run the generator; the
package's own `sdk: ^3.5.0` constraint for consumers is unchanged.
- Fixed the ffigen `-resource-dir` compiler-opt, which pointed at a clang
17 toolchain directory that no longer exists.

## 0.9.0-dev.9 — Gemma-4, MTP removed, dynamic Apple framework

Native rebuild required — `src/llama.cpp` moved `6b4e4bd58` → `d6d0ce82`
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ Then download the platform binary for your project:
| Platform | Artifact | Where to put it |
|---|---|---|
| macOS (dev/test) | `libllama.dylib` + sibling `libggml*.dylib`, `libmtmd.dylib` | anywhere on disk; pass path to `LlamaEngine.spawn` |
| iOS / macOS app | `llama.xcframework` (3 slices: `ios-arm64`, `ios-arm64-simulator`, `macos-arm64`) | drag into Xcode → "Embed & Sign" → call `LlamaEngine.spawnFromProcess` |
| iOS / macOS app | `Llama.xcframework` (3 self-contained slices: `ios-arm64`, `ios-arm64_x86_64-simulator` [universal fat binary], `macos-arm64`) | Auto-linked via Flutter CocoaPods or drag into Xcode → "Embed & Sign" → call `LlamaEngine.spawnFromProcess` |
| Android | `llama-cpp-dart.aar` (CPU + mtmd, arm64-v8a) **or** `llama-cpp-dart-hexagon.aar` (CPU + OpenCL + Hexagon NPU + mtmd, arm64-v8a, Snapdragon) | `android/app/libs/` and `implementation files('libs/llama-cpp-dart.aar')` in Gradle |

Build artifacts yourself with:
Expand Down Expand Up @@ -248,7 +248,7 @@ plan.md // milestone-by-milestone roadmap

`0.9.x` is the rewrite line. The Dart API is mostly stable but **may break once more** before 1.0 — most likely around: real Jinja support, on-device validation findings, and final naming for chat-template/policy knobs. Pin to a minor when you ship.

llama.cpp is pinned per release in `src/llama.cpp` (git submodule). Bumps are tested against the full suite before tagging. The current pin is **tag `b9360`** (sha `6b4e4bd58`); if you're building your own native libs to match this package, check out that tag.
llama.cpp is pinned per release in `src/llama.cpp` (git submodule). Bumps are tested against the full suite before tagging. The current pin is **tag `b10182`** (sha `afeebe103`); if you're building your own native libs to match this package, check out that tag.

## License

Expand Down
58 changes: 58 additions & 0 deletions ios/Llama.xcframework/Info.plist
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>AvailableLibraries</key>
<array>
<dict>
<key>BinaryPath</key>
<string>Llama.framework/Llama</string>
<key>LibraryIdentifier</key>
<string>ios-arm64</string>
<key>LibraryPath</key>
<string>Llama.framework</string>
<key>SupportedArchitectures</key>
<array>
<string>arm64</string>
</array>
<key>SupportedPlatform</key>
<string>ios</string>
</dict>
<dict>
<key>BinaryPath</key>
<string>Llama.framework/Llama</string>
<key>LibraryIdentifier</key>
<string>ios-arm64_x86_64-simulator</string>
<key>LibraryPath</key>
<string>Llama.framework</string>
<key>SupportedArchitectures</key>
<array>
<string>arm64</string>
<string>x86_64</string>
</array>
<key>SupportedPlatform</key>
<string>ios</string>
<key>SupportedPlatformVariant</key>
<string>simulator</string>
</dict>
<dict>
<key>BinaryPath</key>
<string>Llama.framework/Versions/A/Llama</string>
<key>LibraryIdentifier</key>
<string>macos-arm64</string>
<key>LibraryPath</key>
<string>Llama.framework</string>
<key>SupportedArchitectures</key>
<array>
<string>arm64</string>
</array>
<key>SupportedPlatform</key>
<string>macos</string>
</dict>
</array>
<key>CFBundlePackageType</key>
<string>XFWK</string>
<key>XCFrameworkFormatVersion</key>
<string>1.0</string>
</dict>
</plist>
Original file line number Diff line number Diff line change
@@ -0,0 +1,86 @@
#pragma once

#include "ggml.h"

#ifdef __cplusplus
extern "C" {
#endif

typedef struct ggml_backend_buffer_type * ggml_backend_buffer_type_t;
typedef struct ggml_backend_buffer * ggml_backend_buffer_t;
typedef struct ggml_backend * ggml_backend_t;

// Tensor allocator
struct ggml_tallocr {
ggml_backend_buffer_t buffer;
void * base;
size_t alignment;
size_t offset;
};

GGML_API struct ggml_tallocr ggml_tallocr_new(ggml_backend_buffer_t buffer);
GGML_API enum ggml_status ggml_tallocr_alloc(struct ggml_tallocr * talloc, struct ggml_tensor * tensor);

// Graph allocator
/*
Example usage:
ggml_gallocr_t galloc = ggml_gallocr_new(ggml_backend_cpu_buffer_type());

// optional: create a worst-case graph and reserve the buffers to avoid reallocations
ggml_gallocr_reserve(galloc, build_graph(max_batch));

// allocate the graph
struct ggml_cgraph * graph = build_graph(batch);
ggml_gallocr_alloc_graph(galloc, graph);

printf("compute buffer size: %zu bytes\n", ggml_gallocr_get_buffer_size(galloc, 0));

// evaluate the graph
ggml_backend_graph_compute(backend, graph);
*/

// special tensor flags for use with the graph allocator:
// ggml_set_input(): all input tensors are allocated at the beginning of the graph in non-overlapping addresses
// ggml_set_output(): output tensors are never freed and never overwritten

typedef struct ggml_gallocr * ggml_gallocr_t;

GGML_API ggml_gallocr_t ggml_gallocr_new(ggml_backend_buffer_type_t buft);
GGML_API ggml_gallocr_t ggml_gallocr_new_n(ggml_backend_buffer_type_t * bufts, int n_bufs);
GGML_API void ggml_gallocr_free(ggml_gallocr_t galloc);

// pre-allocate buffers from a measure graph - does not allocate or modify the graph
// call with a worst-case graph to avoid buffer reallocations
// not strictly required for single buffer usage: ggml_gallocr_alloc_graph will reallocate the buffers automatically if needed
// returns false if the buffer allocation failed
// ggml_gallocr_resrve_n_size writes the buffer sizes per galloc buffer that would be allocated by ggml_gallocr_reserve_n to sizes
GGML_API bool ggml_gallocr_reserve(ggml_gallocr_t galloc, struct ggml_cgraph * graph);
GGML_API void ggml_gallocr_reserve_n_size(
ggml_gallocr_t galloc,
struct ggml_cgraph * graph,
const int * node_buffer_ids,
const int * leaf_buffer_ids,
size_t * sizes);
GGML_API bool ggml_gallocr_reserve_n(
ggml_gallocr_t galloc,
struct ggml_cgraph * graph,
const int * node_buffer_ids,
const int * leaf_buffer_ids);

// automatic reallocation if the topology changes when using a single buffer
// returns false if using multiple buffers and a re-allocation is needed (call ggml_gallocr_reserve_n first to set the node buffers)
GGML_API bool ggml_gallocr_alloc_graph(ggml_gallocr_t galloc, struct ggml_cgraph * graph);

GGML_API size_t ggml_gallocr_get_buffer_size(ggml_gallocr_t galloc, int buffer_id);

// Utils
// Create a buffer and allocate all the tensors in a ggml_context
// ggml_backend_alloc_ctx_tensors_from_buft_size returns the size of the buffer that would be allocated by ggml_backend_alloc_ctx_tensors_from_buft
// ggml_backend_alloc_ctx_tensors_from_buft returns NULL on failure or if all tensors in ctx are already allocated or zero-sized
GGML_API size_t ggml_backend_alloc_ctx_tensors_from_buft_size(struct ggml_context * ctx, ggml_backend_buffer_type_t buft);
GGML_API struct ggml_backend_buffer * ggml_backend_alloc_ctx_tensors_from_buft(struct ggml_context * ctx, ggml_backend_buffer_type_t buft);
GGML_API struct ggml_backend_buffer * ggml_backend_alloc_ctx_tensors(struct ggml_context * ctx, ggml_backend_t backend);

#ifdef __cplusplus
}
#endif
Loading