StartupXO
Search
English

Posts from pythongiant.github.io

  1. Find XO: KVBoost: Reducing LLM First Token Delay by 5–48× via KV Cache Reuse, Works Without GPU (pythongiant.github.io)

    SummaryThe open-source KVBoost was posted on Hacker News. It splits input into chunks, hashes them, and reuses the KV cache directly when the same chunk reoccurs. It reduced time to first token by 5–48 time…

    1 vote mrlatte Summary Discuss

Keyboard shortcuts

Choose a post with the up and down arrows, then press Enter.

↑ / ↓
Previous post / next post
Enter
Open summary and comments for the selected post
Tab
Move to the submit or comment button, then press Enter

Type normally in text fields. Tab and Enter are always available.