example.com/path/to/article
000 points · username · 0 hours ago
example.com336 points · 199 comments · 13 days ago · raybb
JKCalhoun
amanzi
bambax
For more complex or important tasks, costs, autonomy and privacy matter, but then so does performance/quality.
So I'm not completely convinced it's really worth it; but it's tempting!
akg_67
---
Qwen3.8-27B-4bit, Prompt Processing (PP) 66.3 tok/s, Token Generation (TG) 11.8 tok/s
Ornith-1.5-35B-A3B-MLX-4bit, PP 379.7, TG 45.8
Ornith-1.5-35B-A3B-MLX-4bit, PP 381.5, TG 46.4
Qwen3.6-35B-A3B-mxfp4, PP 389.6, TG 47.6
Qwen3.6-35B-A3B-OptiQ-4bit, PP 342.6, TG 44.4
---
Qwen3.8-27B-4bit generally runs out of output token before completing the task though excellent partial results.
Ornith-1.5-35B-A3B-MLX-4bit seems to get in the loop often specially with tool calls.
Qwen3.6-35B-A3B-mxfp4 seems to be optimal with speed and quality output.
I am going to test Qwen3.6-35B-A3B-4bit soon with same code block just to check my intuition that any derivatives don't seem to perform better than the originals.
jumploops
I've since acquired two DGX Sparks, and it feels so much snappier.
amelius
The main reason to run local: cloud APIs are rented land. They can change their pricing, hit your usage limits, or swap the model being served behind the scenes whenever they feel like it.
Yes but it's easy to replace them.
The main reason should be privacy.
whatsThisBtn4
Meanwhile the stock market has Nvidia at the top... Until everyone gets cuda.
sdevonoes
thrw93747572007
I, personally, do use frontier models in the cloud for a lot of (meta-)cognitive analyses that are heavy enough to have me run against the limits of payed accounts regularly - so I'm neither a Luddite nor stingy with cash in this case.
However: I have pretty good experiences with local models as well. My solid but hardly extreme desktop (with one RX 9070 XT 16GB) mostly serves gemma4:12b and specialized models (embedding) to my local network. This is for general use like simple queries, simple code, reformatting and the like but also for two specific tasks that are permanently running:
a) It's connected to Home Assistant (as a second stage after very simple "turn light XY on" commands which get processed without LLM). So, I can mumble into my smartwatch "computer, how much gas do we have in the warp core and how much energy did the bussard collectors make from the cosmic dust today?" (or describe a more complex light scene or create an automation I want or whatever). The phone transcribes that - with a local model on device - and fires it to the desktop who has agentic access to HA, looks through the sensors and data, sees that I've tagged my solar panels and battery with nerd vocabulary. It makes the right conclusion, converts a few units and gives me back a nice overview. All hands-free while I'm sitting on the toilet.
b) It's the LLM backend for a personal radio station run by a fleet of nerdy/quirky AI DJs who's archetypes are represented more than well enough in the latent space of the "small" model to produce funny results. The DJs can produce consistent, individual segments and programs, run a playlist that works well for me (based on multi-layered audio analysis that also uses local LLMs), respond to song wishes and generally produce much better recommendations than Spotify ever could for me. And you can also put multiple of them in the "studio" to create hilarious crossovers that you would not get from a commercial entity because the IP owners would rather shoot each other in the face.
All of this doesn't even max the available resources, so I can shovel F5-TTS into the VRAM as well and have all my DJs have good, locally created voices (or voice clones of Captain Picard and Han Solo, if I wanted to) based on zero-shot voice cloning.
--> Far from "unusable". It just depends on the task. And I neither have to hand my keys to the Navidrome server nor to my Smart Home to any entity outside my local network.
t1E9mE7JTRjf
My understanding of using a mac mini for ai (ie running a claw bot or whatever) is to have it 'always on' and a better price/performance profile than a cheap vps.
Is there performance (silicon processor?) so unique? As I guess it's not their graphics units. I see tonnes of people using mac minis for AI, to the point it almost became a meme.
Edit: yes I know this article is about local models, my question is a bit more general.
c16
Running a large model locally comes down to one thing: how much RAM it actually needs in memory.
Not completely true. It's memory AND memory bandwidth. You can have 1tb of memory but if you have awful memory-bandwidth you'll also have slow tok/s. A3B helps with this, but so does MTP.
From my experience, you'd be better off running the dense 27b-mlx with MTP than the 3.6 version with A3B. You say your model is ~20GB of ram, but the 3.8:27b-mlx is 18GB and gets me very reasonable tok/s, and greater speed if you disable thinking when not required.
brainless
These are really good models but the harness has to be built around them. I have a ton of generated system prompts for specific purposes. Even parts of a SolidJS stack, for example Route management, has its own prompt. These are experiments but the results are real. If we build harnesses around small models, we can build a locally running WYSIWYG editor which works on plain text prompts.
The performance, in simple tokens/second, is not the most important factor. For many private data points, like emails, I would rather have a local graph based search and LLM on top where the harness is specific to problems like calendar, contacts, finance, etc.
I run all experiments on an 16GB M4 Mac Mini but coding agents building the harness are a mix of Codex, Claude Code and opencode.
apexalpha
Literally today, but it feels like an improvement.
mkagenius
https://x.com/mkagenius/status/2093730391429685732
(xcancel seems to have received a cease and desist)
gigatexal
I wanna get a desktop Mac for local ai so that I don’t turn my laptop into a delta 15k rpm fan when I run things.
I guess I’ll get in line for one hah.
manmal
crossroadsguy
<a href="https://omlx.app">oMLX</a>
Is that supposed to be hallucination? The human or other kind. Feels like a made up URL. It's .ai, isn't it?
ttul
If there was a "Mullvad of GPU clouds", would that solve the privacy concerns?
jimbobthemighty
alexgoodhart
Not many people share setup with actual setup handholding so that was very G of you
stub_out
miles_io
hoistway
thenthenthen
max979
wila
xydac
willtemperley
You do not know what these companies do with your data once they have it. They might limit how it gets used, they might sell it, they might expose it.
This is the burning question for me, what are they doing with our hard work.
I'd have thought that sherlocking a user's $10M business would be too high risk, given the billions at stake if real evidence of this happening was found.
However, OpenAI are currently being sued by Apple for trade secret theft, and the way it was done seems to be abundantly idiotic.
So I'm torn.
mintflow
I am curious is what is the 80% request served by this setup, I was using it for OpenClaw which run serveral cron jobs that discover stuffs over the wide internet, check my support system's unanswered tickets, browser X and some social media for me to filter the valued ones(though I have to say even with GPT 5.6 sol, the quality is low for the timeline X sent to me)
Btw, Tailscale is quite cool and did a good job, I was using it to serve the local LLM and connct the openclaw on a Linux Machine to it.
kelt_row
SipitenoMK
saejox
I read those blog posts to remind of wealth gap i have with the average hackernews user.
Thank you for the encouragement, i will work harder to reach your level.
altern8
I suppose I am waiting for AI-in-a-Box to come along so I can (painlessly) join in.
(I'm sure wrangling with all these esoteric aspects of LLMs though is fun for some people.)