Skip to main content

AMD Halo Station: 96 Cores, 576GB, a Trillion Parameters

AMD Halo Station: 96 Cores, 576GB, a Trillion Parameters

Sept 5, 2026 · by OpenSource Factory · Local AI · Hardware · 7-min read  |  All specs vendor-reported unless marked estimate

I opened the IFA keynote recap expecting another mini-PC refresh. Then Jack Huynh rolled out a liquid-cooled tower with Instinct accelerators inside it and I actually sat up. AMD is calling it the Threadripper Halo Station, and the pitch is simple: datacenter-class memory on your desk, no cloud queue, no shared tenancy.

So here is the short answer up front. This is AMD's DGX Station rival: a 96-core Threadripper PRO plus up to four Instinct MI350P cards, up to 2.6TB of combined memory and up to 16.4TB/s of total bandwidth, aiming at trillion-parameter models run locally. It is a prototype, it ships through OEM partners sometime in 2027, and there is no official price or date yet. Everything below carries that prototype caveat, and I will redo this post the moment real OEM pricing and independent numbers land.

Threadripper Halo Station tower on the IFA 2026 stage
The Halo Station tower on stage at IFA 2026. Big glass case, liquid cooling everywhere, two accelerators visible in the demo unit. (photo: Tom's Hardware)

So what did AMD actually show?

September 4, 2026, Berlin, IFA opening keynote. Huynh, AMD's SVP and GM of Computing and Graphics, framed it as personal AI moving from single prompts to agents that run all day. The Halo Station is the big brother of the Ryzen AI Halo mini-PC box from earlier this year, aimed at the other end of the market: researchers and small teams who are currently renting cloud GPUs by the hour and hating the bill.

AMD's own page calls it a prototype outright and says "Coming in 2027". No partners named yet, no SKU list, no OS confirmed. ServeTheHome's read matches mine: the demo box is assembled from commercial parts, closed-loop radiators per processor rather than a shared loop, an ASUS Pro WS WRX90E-SAGE SE board and an unidentified SFP networking card inside. That is encouraging in one way, since COTS hardware can ship fast, and it also means the final OEM boxes may look different from the stage unit.

AMD CEO on stage next to the Halo Station tower
On stage in Berlin. The tower dwarfs the presenter, and that is kind of the point. (photo: TechPowerUp)

What is inside the tower?

PartHalo Station spec
CPUThreadripper PRO 9995WX, Zen 5 "Shimada Peak", 96 cores / 192 threads, up to 5.4 GHz boost, 384MB L3, 350W
Accelerators2x Instinct MI350P base (288GB HBM3e), path to 4x (576GB). Each card: 144GB HBM3e at 4 TB/s, 128 CDNA4 CUs, up to 600W board power, configurable down to 450W
System memoryUp to 2TB DDR5 RDIMM, 8-channel, up to DDR5-6400
Combined memoryUp to 2.6TB (2TB DDR5 + 576GB HBM3e, 4-card config)
BandwidthUp to 16 TB/s GPU aggregate + up to 410 GB/s RDIMM = up to 16.4 TB/s total
CoolingLiquid cooling on CPU and each accelerator (first liquid-cooled MI350P seen; launch cards were air-cooled for servers)
Platform128x PCIe Gen5 lanes, WRX90 board in demo unit, SFP networking card
SoftwareTo be confirmed. Linux-based stack likely; Windows possible. Microsoft showed its Project Zenith dev toolkit on stage

The CPU is last year's flagship workstation chip, and that is fine. The 9995WX brings 8 channels of DDR5 and 128 PCIe Gen5 lanes, which is exactly what a four-accelerator box needs. The newsworthy part is the MI350P in liquid-cooled form. The card launched in May as a cut-down CDNA4 part for PCIe servers, air-cooled, and a tower with no forced-air server chassis needed something else. Hence the water blocks.

How much memory is that, really?

Here is where I put it next to the boxes I already covered on this blog, using the same numbers as those posts so nothing drifts. The chart below stacks system RAM under accelerator HBM for all four.

Total memory comparison: Halo Station vs DGX Station vs Mac Studio vs DGX Spark
Total usable memory, stacked. The Halo bar is the 4-card config; the IFA demo unit only showed two cards. M5 Ultra 512GB is a build-to-order bump over a 96GB base, same as in my Mac Studio post. (vendor-reported specs, redrawn)

The headline math: Halo at 2,624GB vs DGX Station at 748GB. That divides out to about 3.5x, and AMD claims 3.4x on its page, so the arithmetic roughly checks against their own footnote config. What the chart also shows is the shape of the two designs. Halo piles up a huge 2TB slab of slower DDR5 next to the fast HBM, while DGX Station keeps a smaller 496GB LPDDR5X pool but wires it coherently with the GPU over a 900 GB/s link. Notebookcheck made the same point I would: these are different memory architectures, so raw gigabytes flatter the Halo side. For pure weight capacity the Halo wins by a mile; for how smoothly a model straddling both pools actually runs, coherence matters and AMD has not shown that part yet.

And how fast do tokens come out?

Memory bandwidth comparison across local AI machines
Bandwidth, the number that decides token speed. HBM peaks are per-stack sums and sustained decode runs lower on every box here, but the ordering is what matters. (vendor-reported peaks, redrawn)

I keep repeating this because it is the whole ballgame for local LLMs: every token re-reads the weights, so bandwidth is token speed and capacity is just whether the model fits at all. Halo's 16.4 TB/s tower over DGX Station's ~7.5 TB/s looks absurd until you remember both numbers are sums of HBM-stack peaks, and real decode lives lower. Still, the ordering survives contact with reality: HBM-class boxes play a different sport from unified-memory boxes, and the M5 Ultra's 1.2 TB/s, excellent as it is for a quiet desk box, is an order of magnitude below either station. The honest footnote, and AMD would agree if pressed, is that tensor-parallel chatter across four PCIe cards can bottleneck on the CPU's PCIe fabric. The Register flagged exactly that. Four cards' worth of peak bandwidth only helps if the interconnect keeps up, and we have no independent numbers yet.

Close view of the Halo Station tower interior on stage
Glass off, radiators everywhere. Each processor gets its own closed loop in the demo unit rather than one shared loop. (photo: TechPowerUp)

Can it actually run a trillion-parameter model?

On paper, the FP4 math works. A trillion FP4 parameters weigh roughly half a terabyte, which fits inside 576GB of HBM with room for context and overhead. That is the claim AMD makes, and arithmetically it holds. For the very largest open-weights models, the 2TB DDR5 slab takes the overflow, and The Register name-checked the 2.8-trillion-parameter class as the stretch target. Two things temper this. First, the demo unit AMD showed had two cards, 288GB, so the trillion-parameter story is a four-card story and the four-card box is a "path to", not a SKU. Second, offloaded layers run at DDR5 speed, not HBM speed, so a model that spills is a model that slows. Fitting is not the same as flying.

What will it cost, and what will it drink?

AMD announced no price, and anyone stating one as fact is guessing. The press estimates cluster at $100,000 to $150,000, and Tom's Hardware showed the component-level reasoning: a 9995WX streets around $11,000 to $12,000, each MI350P is estimated near $20,000, and 2TB of DDR5 costs about $50,000 right now in the middle of a RAM shortage. Add chassis, cooling, storage, power and OEM margin and six figures is not a prediction, it is arithmetic. Power is the sleeper issue. Four cards at up to 600W each plus a 350W CPU plus memory and pumps pushes past 3kW at full load by Notebookcheck's estimate, which is server-room territory and possibly a 20-amp circuit conversation in North America. That may be why the stage unit showed two cards. A two-card box sips less, costs far less, and still out-memories everything short of a DGX.

Who is this actually for?

AMD's FAQ names researchers, model developers, simulation users, creators with big generative models, and shops that must keep IP on premises. That reads right to me, with one addition: teams currently paying cloud inference all day. AMD's page says it out loud, keep what you would otherwise spend on cloud inference, and for a lab burning thousands a month on rented HBM, a deskside box starts looking like capital expenditure with a payback schedule rather than a toy. The flip side is the software question, which is genuinely open. Instinct's server stack is real, but a workstation box lives or dies on day-one drivers, ROCm stability on the exact OEM config, and whether Windows is a first-class option. Microsoft's stage cameo suggests Windows matters here, yet nothing is confirmed, and I would not pre-order anything until an OEM puts an OS and a support page next to the price.

Presenter with the Halo Station system on the IFA stage
The pitch is personal AI without the cloud queue. The prototype caveat does the quiet work in that sentence. (photo: TechPowerUp)

My take

As a local-AI-machine story this is the most interesting workstation announcement since the DGX Station itself, and the memory chart is the reason: 576GB of HBM plus 2TB of DDR5 in a tower is a combination nobody else is offering, at any price, and the FP4 trillion-parameter math genuinely fits. The parts that keep me honest are all in the paragraphs above and worth one repeat: prototype, 2027, no price, no partners, demo showed two cards while the big numbers need four, PCIe interconnect unproven at this scale, software unconfirmed, and a power envelope that belongs in a server room with a serious circuit. If the OEM boxes land near $100,000 with a working Linux stack and the four-card config is real, labs will buy these faster than AMD can validate them. If the four-card path stays a path and the price creeps toward the $300,000-class fully-loaded analogy Tom's Hardware cited, it becomes a halo product in the other sense: admired, rarely bought.

I have covered this blog's local-AI ladder from the 2026 value ranking through the M5 Ultra Mac Studio, the Mac mini RAM question and the Xiaomi AI Cube, and the Halo Station slots in at the very top of that ladder: the first box here that even attempts frontier-class weights locally. When OEM pricing, partner names, or an independent benchmark lands, I will rewrite this post the same week, starting with the interconnect question, because that is where paper specs go to be proven or embarrassed.

Sources: AMD Halo Station product page (specs, footnotes HALO-02/SHP-76); ServeTheHome IFA report (COTS build, board, NIC, MI350P background); The Register (price/power estimates, Kimi K3 offload note, PCIe bottleneck flag); Tom's Hardware (component price math, 9995WX street pricing, Lenovo P8 analogy); TechPowerUp (keynote details, 5.4 GHz boost, 600W TBP); Notebookcheck (3kW estimate, coherent-memory caveat, GB300 salvage context). Photos: Tom's Hardware, TechPowerUp. Charts: author's redraws of vendor-reported specs. Related on this blog: 2026 local AI value ranking, M5 Ultra Mac Studio, Qwen3.8 35B-A3B Mac mini, Xiaomi AI Cube.

Comments