Skip to main content

Posts

Showing posts from August, 2026

DeepSeek V4 Flash Vision Exp: Flash Gets Eyes, Now Open Weight

Image
DeepSeek V4 Flash Vision Exp: Flash Gets Eyes, Now Open Weight Aug 31, 2026 · DeepSeek · V4 Flash · Vision · Open Weights  |  All benchmarks vendor-reported, labeled as such DeepSeek shipped nothing but text-only models all year. That changed August 21 on the API — and today it changed again on Hugging Face: deepseek-ai/DeepSeek-V4-Flash-Vision-Exp is now open weight (MIT). I stared at the model card because it's the first V4 that can actually look at an image and still run like Flash. The gist: Same 284B / 13B-active Flash architecture with vision modules added and continued training. On text-only agents it stays on par with Flash 0731 (Terminal Bench 83.9 vs 82.7, DeepSWE 59.3 vs 54.4) — and on multimodal agents it jumps +10.3 on ApexBench to 36.5 , closing to within 2.9 of Opus 4.8. 1M context, MIT, 384 tokens per image at Flash pricing. The image you sent is the Hugging Face social card — deepseek-ai/DeepSeek-V4-Flash-Vision-Exp , 29 likes on day one. API went...

2026 Local AI Machines Ranked by Value: Which Box Actually Earns Its Price?

Image
2026 Local AI Machines Ranked by Value: Which Box Actually Earns Its Price? Aug 31, 2026 · Local AI · Hardware · Mac mini · DGX Spark · 6-min read  |  All prices vendor-reported or street (Danawa) — estimates labeled I priced out every box that claims to run open-weight LLMs locally this year — from a $1,099 Mac mini to a $9,499 Mac Studio — and ran the same math on each: can it load the model at all, and how many tokens per second per $1,000 does it actually give you? The ranking flips hard depending on whether you want a fast 35B-A3B or a 284B monster. TL;DR — My value ranking for 2026: 1. Mac mini M6 24GB ($1,099) — best tok/s per $1,000 for the models most people actually run (Qwen 35B-A3B Q4 at ~56 tok/s = 51 tok/s per $1k). 2. Mac mini M5 Pro 64GB ($2,299) — the speed king for that same tier (Q4 ~102 tok/s, and the only mini that unlocks Q8). 3. Mac Studio M5 Ultra 256GB ($9,499) — the only sane entry to big MoE (DeepSeek V4 Flash 284B, GLM-5.3 ). Expen...

Qwen3.8 35B-A3B on a Mac mini: How Much RAM Do You Actually Need?

Image
Qwen3.8 35B-A3B on a Mac mini: How Much RAM Do You Actually Need? Aug 31, 2026 · Qwen · Mac mini · Local LLM · 5-min read  |  All numbers vendor-reported or author-estimated, labeled as such I almost skipped the Qwen3.8 launch to chase the 27B dense hype. Then a quiet commit in Alibaba's ms-swift repo popped up: Qwen/Qwen3.8-35B-A3B — 35B total, only 3B active. Same MoE recipe that made the 3.6-35B-A3B a darling for 8–16GB VRAM rigs, now on the 3.8 architecture. Not officially announced yet, sitting in the supported-models table. If you're shopping for a Mac mini to run it locally, this is the model you actually want. The short answer: Q4_K_M (~17.5GB) fits a 24GB Mac mini M6 with room for 32K context and does ~56 tok/s. It does not fit a 16GB mini cleanly. If you want Q8 quality (~35GB), you need a 64GB Mac mini M5 Pro — that's where Q8 unlocks ~51 tok/s (and Q4 flies at ~102 tok/s). 32GB M6 is the comfortable headroom pick on the M6 line. I'm writing t...

96% to 11%: GLM-5.3-Flash, Uncensored at the Weight Level

Image
96% to 11%: GLM-5.3-Flash, Uncensored at the Weight Level August 30, 2026 · AI · LLM · Open Weights · GLM · Model Reviews A blood-pressure spike, then a sigh. This is a story I can only tell by first marking the subject line: I'm writing about a release, not teaching you how to use it, and I'm not posting the uncensored version's link to anyone who asked "can it write this for me." Both the teams that shipped it frame it as safety research — whether that framing survives contact with real use is exactly the question worth sitting with. Now, the news: three days after Z.ai officially named GLM-5.3-Flash, the open-weights crowd did what it does to every frontier release — somebody took the guardrails out of the weights and shipped it. The numbers OrcaRouter published are the hook. Refusal behavior collapses from 96%, 93%, 97%, and 93% down to 11%, 12%, 15%, and 18% across the four categories they measured, and XSTest — the rate of wrongly refusing harmless ...

H3 Max: 5 Seconds of Video in Under 3 — and It's Not MiniMax's Model

Image
H3 Max: 5 Seconds of Video in Under 3 — and It's Not MiniMax's Model Open Source Factory · AI Video · Model Reviews · 2026.08.28 I stopped at a number this week: five seconds of video in under three seconds of wall time. Then I read the fine print and stopped again — because the model doing it isn't a shiny new lab release. It's a retrain of MiniMax's open-weights H3, post-trained and served by fal. Here's the short version. fal took H3 — the 33B open model that already topped the video-editing charts — retrained it for prompt adherence and aesthetics, and shipped the result as H3 Max. Independent benchmarks now rank it #1 in image-to-video and #3 in text-to-video (with audio), and it's roughly 35x faster to serve than MiniMax's own endpoint. The catch, if you want to call it one: it tops out at 768p, it's served through fal rather than runnable locally, and the weights aren't out yet. From open weights to a #1-ranked retra...

Nvidia Is Buying the Home of Open Weights in a $12.9B Deal

Image
Nvidia Is Buying the Home of Open Weights in a $12.9B Deal August 27, 2026 · AI · Open Source · Hardware · NVIDIA Image: Hugging Face logo (huggingface.co). I stopped scrolling when I saw this one. The chip company that has made billions selling compute to the AI boom is reportedly buying the place where the models that run on that compute live. Nvidia has agreed — per The Information, citing a source with knowledge of the deal — to buy Hugging Face for $12.9 billion. The home of opensource-machine-learning, where roughly 2.5 million models and 950,000 datasets live, would become part of Nvidia. Before I get into why this matters, the responsible caveat: this is still a "reported" deal, not a signed one. Business Insider, which first reported Hugging Face fielding takeover interest over the weekend, said talks valuing the company at more than $13 billion "had not yet produced a signed agreement and could still atomize." Reuters couldn't independe...

Ox Alpha Was GLM-5.3-Flash, Trained on Chinese Chips Alone

Image
Ox Alpha Was GLM-5.3-Flash, Trained on Chinese Chips Alone August 26, 2026 · AI · LLM · Open Weights · GLM · Model Reviews The anonymous model I wrote about three days ago finally has a name. Z.ai confirmed today that Ox Alpha — the free, week-long, nobody-will-claim-it frontier model on OpenCode — is GLM-5.3-Flash: a 320B-parameter, 18B-active multimodal MoE under MIT license. That part ends a guessing game. The part nobody's putting in the headline is the reveal buried in the announcement: it runs entirely on Chinese AI chips. This is the follow-up I promised when I played detective on Ox Alpha , and honestly the outcome is more interesting than a name reveal. The mystery model wasn't just claimed by Z.ai — it's the first public look at what a frontier-class Chinese open-weights model costs when the training stack has zero NVIDIA hardware in it. The reveal, finally Item GLM-5.3-Flash (Ox Alpha) Parameters 320B total, 18B active (A18B MoE) License MIT Mod...

Qwen3.8-Flash-Next: 6B Active, Qwen4's Architecture Early

Image
Qwen3.8-Flash-Next: 6B Active, Qwen4's Architecture Early August 26, 2026 · AI · LLM · Open Weights · Qwen · Model Reviews Qwen dropped open weights again, but this one isn't another point on the same curve. Qwen3.8-Flash-Next runs 6 billion active parameters out of a 125-billion-parameter model, ships a separate 51-billion-parameter n-gram embedding table you can offload to host RAM, and it's explicitly the preview architecture that the Qwen4 family will be built on. I stopped scrolling when I read that last part — a vendor shipping the next generation's architecture early, on purpose, so the community can hammer on it before Qwen4 lands. Let me be direct about where I land: if you're running local AI on a budget, this is the most interesting model release in a while, and it's not because of a flashy benchmark headline. It's because the architecture was designed with the thing you actually hit — memory — as a first-class constraint. What it actua...

Apple's M5 Ultra Mac Studio: A Real Local AI Machine, at Last

Image
Apple's M5 Ultra Mac Studio: A Real Local AI Machine, at Last August 25, 2026 · AI · Hardware · LLM · Apple Apple finally made the local AI machine I've been waiting for. The new Mac Studio tops out at 512GB of unified memory with 1.2TB/s of bandwidth, and the M5 Ultra is Apple's first quad-die chip. I wrote a skeptical post about a Xiaomi AI Cube prototype last week; this is the opposite story — the memory princess gets an open-sourced software stack to actually use it. Let me be clear about what "real local AI machine" means to me, because it's the whole point of this post. Not a box that runs one vendor's chat app. A machine where I can pull any open model, load it as a GGUF or an MLX build, run, fine-tune, and even cluster. Apple announcing huge memory numbers is one thing; making that memory useful for the open ecosystem is the thing I actually care about. The numbers that justify the price Config CPU / GPU Memory (max) Bandwidth Start...