The openTPU FPGA project, an open-source LLM inference accelerator built largely by AI agents, reached the Hacker News front page this week. Its single Apache-2.0 repository holds the RTL, ISA, compiler and simulator, and it runs Qwen3.5, Gemma 4 and a 35B mixture-of-experts model on a Kintex-7 PCIe card.
What openTPU is and why it went viral
openTPU is a project by GitHub user FeSens. The README frames it around two questions: how far AI agents can go at hardware design, and whether they can build the chip that runs their own inference. According to AICoder’s report, the post collected more than 250 points and 300 comments on Hacker News, and the author said there that the accelerator started at a few tokens per second and reached 80+ tok/s on small models through an iterative self-improvement loop.
What makes it more than a demo is that the whole stack is in one place. The FeSens/openTPU repository contains:
- SystemVerilog RTL for the accelerator, starting at
rtl/top/otpu_top.sv - An instruction set where each instruction is 8 × 32-bit words
- A bit-exact Python ISA simulator that serves as the specification
- A Python-embedded kernel language (
@ol.jit) and its compiler - Host tools (
otpu-chat,otpu-smi) and the Lens profiler, with rooflines and per-cycle timelines
The author also pitches it as a learning project: a repo you can read end to end, “from a matmul in Python down to the wires.”
Inside the openTPU FPGA architecture
The machine is deliberately simple. A sequencer issues one instruction per cycle to a few functional units, and there is no cache and no hidden scheduling. Every data movement is an explicit instruction, so a trace shows exactly where the cycles go.

The units are:
- DMA – moves data between DRAM and on-chip buffers.
- Matrix unit – an int8 systolic array (four columns in the current production image) that multiplies weights streamed from DRAM.
- Vector unit – fp32 math for norms, activations and softmax.
- Quantizer – turns results back into int8 for the next matmul.
The software flow runs top to bottom: kernels written in the ol language, compiled into ISA instructions, executed either by the Python simulator or by the RTL, which is built into a Vivado bitstream for the card. The tests check that the simulator and RTL produce the same bits. A simplified MLP kernel from the README shows the style: ol.load, ol.quantize, ol.dot for the gate, up and down projections, and ol.all_gather to combine results.
The hardware platform
| Item | Value (per the README) |
|---|---|
| Board | Inspur YPCB-00338 PCIe card |
| FPGA | Xilinx Kintex-7 xc7k480t |
| Memory | Two DDR3-1066 channels, 17.1 GB/s peak, 4 GiB |
| Clock | 133.33 MHz, one bitstream for all models |
| Memory controller | LiteDRAM, calibrated by a small CPU inside the memory core (12 s at start-up) |
| Timing margin | WNS +0.032 ns, “only just” closing |
Measured results: what runs and how fast
The README reports ten-plus models running with real weights, with the card producing the same tokens as the simulator, bit for bit. Decode is memory-bound: the card uses 82–94% of DRAM peak while decoding. Selected device-side decode figures (self-reported, measured 29 Sep–1 Oct 2026):
| Model | Weights | Decode (device) | DRAM while decoding |
|---|---|---|---|
| LFM2.5-230M | 4-bit, int8 head | 85.8 tok/s | 14.1 GB/s (82%) |
| Qwen3-0.6B | 4-bit, int8 head | 31.3 tok/s | 13.9 GB/s (82%) |
| Qwen3.5-2B | 4-bit, int8 head | 12.09 tok/s | 15.8 GB/s (92%) |
| Gemma 4 E2B | 4-bit, 4-bit head | 12.14 tok/s | 15.5 GB/s (91%) |
| Phi-4-mini (3.8B) | 4-bit, int8 head | 6.56 tok/s | 15.8 GB/s (92%) |
Two engineering tricks stand out. First, 4-bit weights use FP4 values with two-level block scales (4.25 bits per weight) and keep the LM head in int8; this cuts bytes per token by about a third and raises decode speed by 40–45%, with a perplexity cost documented per model. Second, mixture-of-experts models larger than the 4 GiB card stream their experts from the host. Qwen3.5-35B-A3B (34.7B parameters, 3.0B active) runs at 3.95 tok/s with 62% of expert uses hitting on-card slots and 153 MB streamed per token over PCIe, as described in the offload docs. LFM2.5-8B-A1B reaches 10.6 tok/s.
Keep the caveat in view: AICoder notes these numbers are self-reported, with no independent reproduction or like-for-like GPU comparison yet.
What “AI-designed accelerator” actually means
“Designed by AI” does not mean a model produced a chip in one prompt. openTPU reuses the method from the author’s earlier auto-arch-tournament project, an autonomous loop pointed at a SystemVerilog RV32IM CPU. In that loop, one agent proposes a microarchitectural hypothesis, another implements it in an isolated git worktree, and a fixed evaluation pipeline (Verilator lint, Yosys synthesis, riscv-formal, cosimulation against a Python ISS, three-seed place-and-route plus CoreMark) decides whether it beats the current champion. Of 73 hypotheses in its published run, 63 were rejected by the verifier.
The author’s own conclusion is the lesson worth keeping: the loop is commodity, and the verifier is the moat. openTPU applies the same thinking. The bit-exact simulator is the spec, every RTL or ISA change must keep pytest passing, the card is checked token-for-token against the simulator, and a tools/validate.py script compares outputs with a Hugging Face golden model using top-1 agreement and KL divergence. The README also mentions “a tournament of Vivado runs” still working on timing margin and area. Agents generate; strict, automated checks decide what survives. That is the same pattern we see in where AI circuit design wins versus traditional engineering.
FPGA lessons for embedded engineers
Whether or not you care about LLMs, the openTPU FPGA design is a compact case study in practical accelerator engineering:
- Find the real bottleneck. Decode is bound by DRAM, not compute. A faster clock mostly helps prefill, so the work focuses on squeezing the last few percent of DDR3 efficiency.
- Match the clock to the memory. 133.33 MHz was chosen because it is where the 128-byte port matches the two DDR3 channels.
- Quantize where the bytes are. Cutting weights to ~4.25 bits per weight directly buys speed in a memory-bound design.
- Make the hardware predictable. No cache, no hidden scheduling, one instruction per cycle: simpler to verify, simpler to profile.
- Treat a golden model as the spec. A bit-exact simulator lets most development happen without hardware, and makes board bugs obvious.
- Software decides adoption. A compiler, profiler and chat tool make the hardware usable, a point we made in why edge AI accelerator success depends on its software tools.
How to try openTPU without an FPGA
According to the README, everything except the card runs on a laptop. The steps are:
- Clone the repo and install it:
pip install -e ., thenpip install pytest torch transformers. - Run the test suite:
python3 -m pytest -q(the RTL tests also need Verilator 5). - Download a small model:
hf download LiquidAI/LFM2.5-230M --local-dir models/LFM2.5-230M. - Chat on the simulator:
otpu-chat --model lfm2 --backend isa. - Read in this order:
docs/isa.md, the kernels anddocs/compiler.md,opentpu/isasim.py, then the RTL.
With the card, you build the bitstream with make bit in boards/ypcb-00338, load it over JTAG, run sudo otpu-setup and then otpu-chat --backend board. Contributions are welcome, and the author notes most of the work needs only Python and Verilator. If you are wondering whether FPGA skills are still worth building next to microcontrollers, our take on whether FPGA is old technology is a good companion read.
Key takeaways
- openTPU is an Apache-2.0, AI-agent-developed LLM accelerator with RTL, ISA, simulator, compiler and profiler in one repo.
- On a Kintex-7 xc7k480t card it reports up to 85.8 tok/s (LFM2.5-230M, 4-bit) and 3.95 tok/s on Qwen3.5-35B-A3B with host-streamed experts.
- Decode is DRAM-bound; 4-bit weights give a 40–45% decode speed-up.
- The AI-designed accelerator story is really a verification story: bit-exact simulators and automated gates make agent-generated RTL trustworthy.
- Figures are self-reported and not yet independently reproduced.
FAQ
What is openTPU?
openTPU is an open-source (Apache-2.0) AI inference accelerator developed largely by AI agents. Its repository includes SystemVerilog RTL, an ISA, a bit-exact Python simulator, a kernel compiler and a profiler, and it runs LLMs on a Kintex-7 FPGA PCIe card.
Which models run on the openTPU FPGA card?
The README lists LFM2.5, LFM2, Qwen3, Qwen3.5 (0.8B, 2B, 4B), Gemma 4 E2B and E4B, SmolLM3-3B and Phi-4-mini, plus MoE models LFM2.5-8B-A1B and Qwen3.5-35B-A3B with experts streamed from the host.
Do I need an FPGA to use openTPU?
No. Everything except the card runs on a laptop through the Python ISA simulator, and RTL tests run in Verilator 5. You only need the Inspur YPCB-00338 card to run on real hardware.
Was openTPU entirely designed by AI?
The project describes itself as “developed by AI,” using an agent loop similar to auto-arch-tournament, but a human author built the verification framework and evaluation pipeline that decide which changes are accepted. Treat it as AI-driven design under human-defined checks.
Want to build the FPGA, embedded AI and hardware design skills behind projects like this? Explore the Educational Engineering Team courses at eduengteam.com/.
