Search results for

“MLX”

76 results found

Post

Replies

Boosts

Views

Activity

MLX,MLX LM, MLX LM Server -> Is there a bootstrap repo?
theres a MLX, a MLX LM and a MLX LM Server mentioned. Is there a Bootstrap GitHub repo out there that can be used to directly, and quickly, set up an example of this, without the hassle of setting up, kind of like a bootstrap for us mere mortals? And what is the feasibility of using these on a M3 Pro with 18Gb of memory? - can these be bounced between a local M3 Pro and a Tailscale-linked M2 Pro with 36Gb memory? Do both need to be on macOS27 for it to work?
1
0
139
Jun ’26
Data used for MLX fine-tuning
The WWDC25: Explore large language models on Apple silicon with MLX video talks about using your own data to fine-tune a large language model. But the video doesn't explain what kind of data can be used. The video just shows the command to use and how to point to the data folder. Can I use PDFs, Word documents, Markdown files to train the model? Are there any code examples on GitHub that demonstrate how to do this?
2
0
722
Jun ’25
Fused Metal Kernels for Linear Recurrences in MLX
I’ve been developing mlx-recurrence, a plug-in framework of fused Metal GPU kernels for linear recurrences on Apple silicon—roughly analogous to flash linear attention for MLX. Sequential recurrences are difficult for MLX to fuse automatically. Architectures such as state-space models, gated linear attention, and diagonal RNNs ordinarily require a loop across the sequence length. When that loop is implemented in Python, a sequence of length L can require L separate Python-to-Metal dispatches. These kernels instead execute the entire recurrence in a single Metal dispatch. The training path uses segment checkpointing with recomputation during the backward pass. In validated M3 Max tests, the checkpoint-and-recompute kernels reduced peak recurrent-state memory by approximately 12–18× at the kernel level and lowered total training peak memory from 23.88 GB to 10.34 GB. At the same batch size, end-to-end training throughput improved by roughly 1.4×, while individual fused forward-and-bac
0
0
169
Aug ’26
Sharing a Swift port of Gemma 4 for mlx-swift-lm — feedback welcome
Hi all, I've been working on a pure-Swift port of Google's Gemma 4 text decoder that plugs into mlx-swift-lm as a sidecar model registration. Sharing it here in case anyone else hit the same wall I did, and to get feedback from the MLX team and the community before I propose anything upstream. Repo: https://github.com/yejingyang8963-byte/Swift-gemma4-core Why As of mlx-swift-lm 2.31.x, Gemma 4 isn't supported out of the box. The obvious workaround — reusing the Gemma 3 text implementation with a patched config — fails at weight load because Gemma 4 differs from Gemma 3 in several structural places. The chat-template path through swift-jinja 1.x also silently corrupts the prompt, so the model loads but generates incoherent text. What's in the package A from-scratch Swift implementation of the Gemma 4 decoder (Configuration, Layers, Attention, MLP, RoPE, DecoderLayer) Per-Layer Embedding (PLE) support — the shared embedding table that feeds every decoder layer through a gated MLP as a
1
0
815
Apr ’26
MLX/Ollama Benchmarking Suite - Open Source and Free
Hi all, I spent the last few months developing an MLX/Ollama local AI Benchmarking suite for Apple Silicon, written in pure Swift and signed with an Apple Developer Certificate, open source, GPL, and free. I would love some feedback to continue development. It is the only benchmarking suite I know of that supports live power metrics and MLX natively, as well as quick exports for benchmark results, and an arena mode, Model A vs B with history. I really want this project to succeed, and have widespread use, so getting 75 stars on the github repo makes it eligible for Homebrew/Cask distribution. Github Repo
0
0
446
Feb ’26
Local Agentic AI on Mac using MLX: issues solved with Gemma-4
I was keen on trying local models with Xcode agents after watching the WWDC 2026 session Run Local agentic AI on the Mac using MLX https://developer.apple.com/videos/play/wwdc2026/232/ Ran into a few issues while following the three setup steps shown in the session, so I put together a small project with the workarounds I used: https://github.com/jdhark-com/opencode_mlx_bridge/ I needed to use another model than the one demonstrated. The main issue I hit was that running: mlx_lm.server --model mlx-community/gemma-4-e4b-it-4bit failed with: ValueError: Received 126 parameters not in model The workaround in start_xcode_server.py is to load the model with strict=False which resolved the issue for me. opencode.json prompt config really helped to get more verbose feedback from the model. Hopefully this helps anyone else trying to get a local MLX model working as an Xcode agent.
0
0
223
Jun ’26
LLM size for fine-tuning using MLX in MacBook
Hi, recently i tried to fine-tune Gemma-2-2b mlx model on my macbook (24 GB UMA). The code started running, after few seconds i saw swap size reaching 50GB and ram around 23 GB and then it stopped. I ran the Gemma-2-2b (cuda) on colab, it ran and occupied 27 GB on A100 gpu and worked fine. Here i didn't experienced swap issue. Now my question is if my UMA was more than 27 GB, i also would not have experienced swap disk issue. Thanks.
1
0
532
Oct ’25
How can a local AI agent use MLX/Metal unattended on macOS while remaining confined to an authorized workspace?
How can a local AI agent use MLX/Metal unattended while remaining confined to an authorized workspace? I am developing an AI-driven local media-processing workflow on an Apple-silicon Mac and am trying to understand the correct architecture for allowing it to run unattended without giving the AI agent unrestricted access to my primary personal computer. I am not a software engineer, so I may be missing an established macOS mechanism or using the wrong terminology. I would appreciate guidance from people familiar with MLX, Metal, sandboxing, and macOS security. What I am building I use OpenAI Codex as the local execution/software-development agent. The working system currently: ingests and verifies original video and still media while preserving immutable originals; performs visual semantic analysis and divides video into meaningful time-coded segments; separately analyzes spoken language rather than assuming audio and video are semantically equivalent; uses MLX Whisper locally on Ap
0
0
371
2w
MLX,MLX LM, MLX LM Server -> Is there a bootstrap repo?
theres a MLX, a MLX LM and a MLX LM Server mentioned. Is there a Bootstrap GitHub repo out there that can be used to directly, and quickly, set up an example of this, without the hassle of setting up, kind of like a bootstrap for us mere mortals? And what is the feasibility of using these on a M3 Pro with 18Gb of memory? - can these be bounced between a local M3 Pro and a Tailscale-linked M2 Pro with 36Gb memory? Do both need to be on macOS27 for it to work?
Replies
1
Boosts
0
Views
139
Activity
Jun ’26
MLX support on swift playground
i cant use mlx on swift for some reason, i would like for them to add the support to add it as a package
Replies
2
Boosts
0
Views
1.1k
Activity
4w
Data used for MLX fine-tuning
The WWDC25: Explore large language models on Apple silicon with MLX video talks about using your own data to fine-tune a large language model. But the video doesn't explain what kind of data can be used. The video just shows the command to use and how to point to the data folder. Can I use PDFs, Word documents, Markdown files to train the model? Are there any code examples on GitHub that demonstrate how to do this?
Replies
2
Boosts
0
Views
722
Activity
Jun ’25
Is there an easy way to convert a MLX format model to Core ML
Hi, I'd like to use a MLX model already in the MLX Community in my App. I understand I first need to convert it to Core ML format. Is there an easy way to do that considering MLX is an Apple project? It would be good if it was easier then I'd be more motivated to use MLX to train my models. Thanks Richard
Replies
1
Boosts
0
Views
2.0k
Activity
Jun ’24
Integrating MLX Models with React Native for iOS Deployment
Hi, I'm looking for the best way to use MLX models, particularly those I've fine-tuned, within a React Native application on iOS devices. Is there a recommended integration path or specific API for bridging MLX's capabilities to React Native for deployment on iPhones and iPads?
Replies
1
Boosts
0
Views
321
Activity
Jun ’25
Fused Metal Kernels for Linear Recurrences in MLX
I’ve been developing mlx-recurrence, a plug-in framework of fused Metal GPU kernels for linear recurrences on Apple silicon—roughly analogous to flash linear attention for MLX. Sequential recurrences are difficult for MLX to fuse automatically. Architectures such as state-space models, gated linear attention, and diagonal RNNs ordinarily require a loop across the sequence length. When that loop is implemented in Python, a sequence of length L can require L separate Python-to-Metal dispatches. These kernels instead execute the entire recurrence in a single Metal dispatch. The training path uses segment checkpointing with recomputation during the backward pass. In validated M3 Max tests, the checkpoint-and-recompute kernels reduced peak recurrent-state memory by approximately 12–18× at the kernel level and lowered total training peak memory from 23.88 GB to 10.34 GB. At the same batch size, end-to-end training throughput improved by roughly 1.4×, while individual fused forward-and-bac
Replies
0
Boosts
0
Views
169
Activity
Aug ’26
Sharing a Swift port of Gemma 4 for mlx-swift-lm — feedback welcome
Hi all, I've been working on a pure-Swift port of Google's Gemma 4 text decoder that plugs into mlx-swift-lm as a sidecar model registration. Sharing it here in case anyone else hit the same wall I did, and to get feedback from the MLX team and the community before I propose anything upstream. Repo: https://github.com/yejingyang8963-byte/Swift-gemma4-core Why As of mlx-swift-lm 2.31.x, Gemma 4 isn't supported out of the box. The obvious workaround — reusing the Gemma 3 text implementation with a patched config — fails at weight load because Gemma 4 differs from Gemma 3 in several structural places. The chat-template path through swift-jinja 1.x also silently corrupts the prompt, so the model loads but generates incoherent text. What's in the package A from-scratch Swift implementation of the Gemma 4 decoder (Configuration, Layers, Attention, MLP, RoPE, DecoderLayer) Per-Layer Embedding (PLE) support — the shared embedding table that feeds every decoder layer through a gated MLP as a
Replies
1
Boosts
0
Views
815
Activity
Apr ’26
MLX/Ollama Benchmarking Suite - Open Source and Free
Hi all, I spent the last few months developing an MLX/Ollama local AI Benchmarking suite for Apple Silicon, written in pure Swift and signed with an Apple Developer Certificate, open source, GPL, and free. I would love some feedback to continue development. It is the only benchmarking suite I know of that supports live power metrics and MLX natively, as well as quick exports for benchmark results, and an arena mode, Model A vs B with history. I really want this project to succeed, and have widespread use, so getting 75 stars on the github repo makes it eligible for Homebrew/Cask distribution. Github Repo
Replies
0
Boosts
0
Views
446
Activity
Feb ’26
Local Agentic AI on Mac using MLX: issues solved with Gemma-4
I was keen on trying local models with Xcode agents after watching the WWDC 2026 session Run Local agentic AI on the Mac using MLX https://developer.apple.com/videos/play/wwdc2026/232/ Ran into a few issues while following the three setup steps shown in the session, so I put together a small project with the workarounds I used: https://github.com/jdhark-com/opencode_mlx_bridge/ I needed to use another model than the one demonstrated. The main issue I hit was that running: mlx_lm.server --model mlx-community/gemma-4-e4b-it-4bit failed with: ValueError: Received 126 parameters not in model The workaround in start_xcode_server.py is to load the model with strict=False which resolved the issue for me. opencode.json prompt config really helped to get more verbose feedback from the model. Hopefully this helps anyone else trying to get a local MLX model working as an Xcode agent.
Replies
0
Boosts
0
Views
223
Activity
Jun ’26
Reply to Bring an LLM provider to the Foundation Models, missing MLX dependencies
This is being introduced to mlx-swift-lm in PR#334 (see here: https://github.com/ml-explore/mlx-swift-lm/pull/334).
Replies
Boosts
Views
Activity
Jun ’26
Reply to Tensor Flow Metal 1.2.0 on M2 Fails to converge on common toy models
@txoof fair enough have you had a go with MLX ? here is some CIFAR-10 example on GitHub https://github.com/ml-explore/mlx-examples/blob/main/cifar/README.md
Topic: Machine Learning & AI SubTopic: Core ML Tags:
Replies
Boosts
Views
Activity
Mar ’25
Reply to MLX,MLX LM, MLX LM Server -> Is there a bootstrap repo?
I would suggest heading over to https://github.com/ml-explore/mlx-swift-lm to see if that package has what you're looking for. They may be able to help more over there for your MLX memory-related questions. :)
Replies
Boosts
Views
Activity
Jun ’26
Reply to LLM size for fine-tuning using MLX in MacBook
Hello, The MLX folks are requesting that you create an issue in the MLX GitHub repo with steps to reproduce the problem. They aim to debug the problem. We'd greatly appreciate it if you posted any solutions back here to the developer forums.
Replies
Boosts
Views
Activity
Oct ’25
LLM size for fine-tuning using MLX in MacBook
Hi, recently i tried to fine-tune Gemma-2-2b mlx model on my macbook (24 GB UMA). The code started running, after few seconds i saw swap size reaching 50GB and ram around 23 GB and then it stopped. I ran the Gemma-2-2b (cuda) on colab, it ran and occupied 27 GB on A100 gpu and worked fine. Here i didn't experienced swap issue. Now my question is if my UMA was more than 27 GB, i also would not have experienced swap disk issue. Thanks.
Replies
1
Boosts
0
Views
532
Activity
Oct ’25
How can a local AI agent use MLX/Metal unattended on macOS while remaining confined to an authorized workspace?
How can a local AI agent use MLX/Metal unattended while remaining confined to an authorized workspace? I am developing an AI-driven local media-processing workflow on an Apple-silicon Mac and am trying to understand the correct architecture for allowing it to run unattended without giving the AI agent unrestricted access to my primary personal computer. I am not a software engineer, so I may be missing an established macOS mechanism or using the wrong terminology. I would appreciate guidance from people familiar with MLX, Metal, sandboxing, and macOS security. What I am building I use OpenAI Codex as the local execution/software-development agent. The working system currently: ingests and verifies original video and still media while preserving immutable originals; performs visual semantic analysis and divides video into meaningful time-coded segments; separately analyzes spoken language rather than assuming audio and video are semantically equivalent; uses MLX Whisper locally on Ap
Replies
0
Boosts
0
Views
371
Activity
2w