OpenTPU is an open-source AI accelerator developed by AI and published by FeSens on GitHub. The repository README lists RTL, an ISA, a simulator, a compiler, and a profiler. The GitHub page captured in the source pool shows 333 stars and 19 forks. El Ecosistema Startup’s October 6, 2026 article describes it as an AI-designed open-source inference chip.
The GitHub description names Qwen3, LFM2.5, and Qwen3.5 running on a Kintex-7 PCIe card. The project reports that FPGA results match the simulator; the README results table also publishes measurements for named model configurations.
TL;DR: OpenTPU is an AI-developed open-source accelerator from FeSens. The GitHub source lists 333 stars and 19 forks. Its README groups RTL, an ISA, a simulator, a compiler, and a profiler in one repository, says it runs ten modern models, and reports FPGA results that match simulation. The repository description names Qwen3, LFM2.5, and Qwen3.5 on a Kintex-7 PCIe card.
What Is OpenTPU and Who Built It?
OpenTPU is an open-source AI accelerator published on GitHub under the FeSens organization. Its README says it was developed by AI and lists RTL, an ISA, a simulator, a compiler, and a profiler. The repository also describes itself as a small monorepo and learning project, intended to be read end to end.
The GitHub excerpt lists 333 stars and 19 forks. El Ecosistema Startup’s Spanish-language article, published on October 6, 2026, describes openTPU as an AI-designed open-source inference chip that runs on an FPGA.
The README says OpenTPU brings lessons from the earlier FeSens repository, auto-arch-tournament, to AI accelerators. That link is the project’s own reference to the related work.
The project materials describe OpenTPU as developed by AI and identify FeSens as its publisher.
How Can an AI Design Its Own Hardware Accelerator?
The README links OpenTPU to auto-arch-tournament and says the accelerator project brings lessons from that repository to hardware design. It also poses two questions: how far AI agents can go in hardware design, and whether they can build the chip that runs their own inference.
The README lists RTL, an instruction set architecture, a bit-exact simulator, a kernel language and compiler, and host software for a real PCIe card. The GitHub description also names a profiler. The README’s results section identifies an Inspur YPCB-00338 board with a Xilinx Kintex-7 xc7k480t and two DDR3 channels.
The repository description names Qwen3, LFM2.5, and Qwen3.5 on a Kintex-7 PCIe card. The README says ten modern models run with real weights and reports that the card produces the same tokens as the simulator, bit for bit. It also provides benchmark and memory-bandwidth tables.
OpenTPU’s materials describe an AI-developed accelerator and report model runs on an FPGA. They document this project and its reported results, not a general conclusion about what AI systems can design.
What’s Inside the OpenTPU Repository?
One repository, five components. The GitHub description is explicit about the contents, and each piece plays a distinct role in the accelerator stack:
- RTL — the register-transfer-level hardware description, the actual design that gets synthesized onto the FPGA
- ISA — the instruction set architecture, the contract between compiled software and hardware execution
- Simulator — a software model of the accelerator that predicts behavior before hardware runs
- Compiler — the toolchain that transforms model code into instructions for the ISA
- Profiler — the measurement layer that shows where execution time goes
The README calls OpenTPU a small monorepo and learning project that can be read end to end. It lists the hardware design, instruction set, simulator, kernel language and compiler, and host software in the project repository.
The GitHub excerpt lists 333 stars and 19 forks. These are repository activity counts captured in the supplied source material.
The README identifies auto-arch-tournament as an earlier FeSens repository and says OpenTPU brings its lessons to AI accelerators. Both repositories are public and linked from the project materials.
The project is described as open source, and its public repository lists the RTL, simulator, and compiler for readers to inspect. Its README also presents the repository as a learning project.
Which Models Can OpenTPU Actually Run?
The short repository description names Qwen3, LFM2.5, and Qwen3.5. The README says OpenTPU runs ten modern models with real weights and provides a table of results for named configurations.
The source excerpts identify these model names but do not explain how model architectures map to the compiler or ISA. The README’s statement that ten modern models run on OpenTPU is the broader support claim; its results table should be read for the configurations it actually reports.
The repository description names a Kintex-7 PCIe card. The README gives a more specific target: an Inspur YPCB-00338 board with a Xilinx Kintex-7 xc7k480t and two DDR3 channels. Those are project-reported hardware details for its documented runs.
The README says the FPGA card produces the same tokens as the simulator, bit for bit, for the project’s reported runs.
The repository description names three models; the README says ten modern models run on OpenTPU and its results table lists specific model configurations.
How Does the FPGA Implementation Compare to the Simulator?
The El Ecosistema Startup article describes results on the Kintex-7 PCIe card as “con resultados idénticos a los del simulador” — identical to the simulator’s results. The repository README also describes bit-for-bit matching for its reported runs.
The README reports that the card produces the same tokens as the simulator, bit for bit. Its results table also publishes decode, prefill, and DRAM-bandwidth measurements for the listed model configurations.
OpenTPU’s repository description names Qwen3, LFM2.5, and Qwen3.5. The README says ten modern models run on the design and reports matching FPGA and simulator outputs for its runs.
The README performance table publishes performance and memory-bandwidth figures, so it would be incorrect to say that no throughput data is available. Its int8 table reports LFM2.5-230M at 59.0 device decode tokens/s, 52.3 wall decode tokens/s, 295.6 prefill tokens/s, and 14.5 GB/s DRAM bandwidth (85% of peak). It reports Qwen3-0.6B at 21.6/21.3/92.1 tokens/s and 14.4 GB/s (84%), and Qwen3.5-0.8B at 17.6/16.3/61.4 tokens/s and 14.5 GB/s (85%). These are project-reported measurements; they do not by themselves establish a performance bottleneck.
The public repository lists the RTL, ISA, compiler, and simulator. The README describes OpenTPU as a learning project and says the monorepo can be read end to end.
What Is auto-arch-tournament and Why Does It Matter Here?
The README says OpenTPU “brings the lessons of auto-arch-tournament to AI accelerators” and links to the separate FeSens repository. It poses the questions of how far AI agents can go in hardware design and whether they can build the chip that runs their own inference.
The README frames OpenTPU around two questions: how far AI agents can go in hardware design, and whether they can build the chip that runs their own inference.
The README links OpenTPU to auto-arch-tournament and says it brings that project’s lessons to AI accelerators. The two public repositories are presented together as related FeSens projects.
How Mature Is the Project on GitHub?
The GitHub excerpt shows 333 stars and 19 forks. They are snapshot counts of activity on the repository page.
The sources describe OpenTPU as an AI-developed open-source accelerator, and the README calls it a learning project. The project materials report FPGA runs with outputs matching the simulator.
The README lists RTL, an ISA, a simulator, a compiler, and a profiler in one repository. It calls OpenTPU a small monorepo and learning project that can be read end to end.
What Does OpenTPU Mean for Open AI Hardware?
OpenTPU’s README lists RTL, an ISA, a simulator, a compiler, and a profiler. The project materials describe those pieces as one AI-developed accelerator stack.
The sources describe OpenTPU as developed by AI and report model runs on an FPGA. The README says the project asks how far AI agents can go in hardware design and whether they can build the chip that runs their own inference.
The README reports that the FPGA card produces the same tokens as the simulator, bit for bit, and includes performance results for named model configurations.
How Can You Get Started With OpenTPU?
The entry point is the FeSens/openTPU repository on GitHub. The README states the project’s premise up front: an open-source AI accelerator developed by AI, bringing the lessons of auto-arch-tournament to accelerator design. Everything needed to engage with the project lives in that single repository.
A practical starting sequence looks like this:
- Clone or browse the FeSens/openTPU repository on GitHub.
- Follow the README’s link to the public auto-arch-tournament repository, which it identifies as related work.
- Study the ISA specification to understand the instruction set the accelerator executes.
- Use the included simulator to run models without any hardware at all.
- Examine the compiler to see how model graphs map onto the ISA.
- Review the profiler output to understand where execution time goes.
- Only then consider the RTL and the Kintex-7 PCIe card target for physical runs.
- The README reports bit-for-bit matching between the card’s outputs and simulator results for the documented runs.
The README describes OpenTPU as a learning project and says the monorepo can be read end to end. It lists a simulator, compiler, profiler, and RTL among the project’s components.
What Are the Limits of the Project Today?
The README identifies a Kintex-7 PCIe target and calls OpenTPU an AI-developed learning project. Its results section says ten modern models run with real weights and includes benchmark and DRAM-bandwidth measurements.
The repository description names a Kintex-7 PCIe card and three models. The README says ten modern models run on OpenTPU and includes performance tables for named configurations on an Inspur YPCB-00338 board with a Xilinx Kintex-7 xc7k480t and two DDR3 channels.
The README reports that the card produces the same tokens as the simulator, bit for bit, for its ten-model project results. The table lists the specific model configurations and measurements included in the documentation.
Frequently Asked Questions
Which models run on OpenTPU?
The repository description names Qwen3, LFM2.5, and Qwen3.5 on a Kintex-7 PCIe card. Its README says ten modern models run on OpenTPU and reports matching FPGA and simulator results for documented runs.
Is OpenTPU free to use?
OpenTPU is published under the Apache 2.0 license, according to El Ecosistema Startup. Its GitHub repository is public and described as open source.
What hardware does OpenTPU target?
The repository description names a Kintex-7 PCIe card; the README specifies an Inspur YPCB-00338 board with a Xilinx Kintex-7 xc7k480t and two DDR3 channels. The sources list a simulator and compiler but do not prescribe a particular workflow.
What performance data does the README include?
Its int8 results table reports decode, prefill, and DRAM-bandwidth measurements for named model configurations, including LFM2.5-230M, Qwen3-0.6B, and Qwen3.5-0.8B. These are figures reported by the project.
Summary
OpenTPU’s README describes a learning project that brings hardware, software, and reported FPGA results into one repository. Key details:
A listed hardware/software stack. The README names RTL, an ISA, a simulator, a compiler, and a profiler in the FeSens/openTPU repository.
AI-developed project. The sources describe OpenTPU as developed by AI and say it brings lessons from the auto-arch-tournament repository to AI accelerators; the provided excerpts do not detail the development workflow.
Reported model runs. The repository description names Qwen3, LFM2.5, and Qwen3.5 on a Kintex-7 PCIe card; the README says ten modern models run on the design and reports FPGA results matching simulation.
Repository activity counts. The GitHub excerpt lists 333 stars and 19 forks. These counts alone do not establish team size, support capacity, or project maturity.
If open accelerator design interests you, clone the repository, run a model in the simulator, and form your own view of what AI-designed hardware can deliver.