kev-gpt
I just posted a Show HN of my most recent side project, a live demo of a tiny-llm implemented in FPGA fabric, hitting an aggregate peak of 60,000tok/s, but a 'usable' model at 21,000tok/s Writeup and demo here: https://www.mikeayles.com/blog/on-chip-llm-kv260/ Source and HDL here: https://github.com/MichaelAyles/kev-gpt Show HN here: https://news.ycombinator.com/item?id=49242475