mirror of https://github.com/ollama/ollama.git synced 2026-02-19 15:57:07 -05:00

Files

Patrick Devine a0407d07fa safetensors quantization for mlx (#14184 )

This change includes:
  - changes to the safetensors metadata format
  - changes to the create command to properly create the blobs with the new format
  - changes to load the new format
  - fixes ollama show to properly show each tensor

2026-02-10 11:29:17 -08:00

CMakeLists.txt

Revert "move tokenizers to separate package (#13825 )" (#14111 )

2026-02-05 20:49:08 -08:00

compile.go

MLX - dynamic loading of mlx-c (#13735 )

2026-01-16 16:34:22 -08:00

doc.go

MLX - dynamic loading of mlx-c (#13735 )

2026-01-16 16:34:22 -08:00

generate_wrappers.go

MLX - dynamic loading of mlx-c (#13735 )

2026-01-16 16:34:22 -08:00

mlx_dynamic.c

MLX - dynamic loading of mlx-c (#13735 )

2026-01-16 16:34:22 -08:00

mlx_dynamic.h

MLX - dynamic loading of mlx-c (#13735 )

2026-01-16 16:34:22 -08:00

mlx_test.go

MLX - dynamic loading of mlx-c (#13735 )

2026-01-16 16:34:22 -08:00

mlx.c

MLX - dynamic loading of mlx-c (#13735 )

2026-01-16 16:34:22 -08:00

mlx.go

safetensors quantization for mlx (#14184 )

2026-02-10 11:29:17 -08:00

mlx.h

MLX - dynamic loading of mlx-c (#13735 )

2026-01-16 16:34:22 -08:00

README.md

Add experimental MLX backend and engine with imagegen support (#13648 )

2026-01-08 16:18:59 -08:00

README.md

MLX Memory Management

| This package will get consolidated with x/ml/backend/mlx in the future.

Automatic Tracking

All arrays are automatically tracked when created. On Eval(), non-kept arrays are freed.

API

result := mlx.Matmul(x, w) // arrays automatically tracked
mlx.Eval(result)           // free non-kept, eval result (auto-kept)

Key Functions

mlx.Eval(outputs...) - free non-kept arrays, then evaluate (outputs auto-kept)
mlx.AsyncEval(outputs...) - async version of Eval (outputs auto-kept)
mlx.Keep(arrays...) - mark arrays to survive cleanup (for weights, caches)
array.Free() - mark array for cleanup on next Eval

Loop Pattern

for step := 0; step < maxTokens; step++ {
    logits := model.Forward(token, caches)
    oldToken := token
    token = sample(logits)

    // Keep cache state across iterations
    for _, c := range caches {
        mlx.Keep(c.State()...)
    }

    oldToken.Free()       // mark for cleanup
    mlx.AsyncEval(token)  // frees old, evals new
}

Notes

Eval() and AsyncEval() auto-keep their outputs
Free() marks for cleanup - actual free happens during next Eval
Use Keep() for weights and cache state that must survive multiple Eval cycles
Arrays created inside compiled closures are managed by MLX, not tracked