Smaller, faster, safer: running Kimi and GLM at scale
blog.cloudflare.com - 73 poäng - 17 kommentarer - 14044 sekunder sedan
Kommentarer (3)
- scrlk - 6240 sekunder sedanNice to see a provider being transparent about KV cache quantisation. I've been suspecting that some providers do this silently whilst heavily promoting their unquantised weights, even though KV quantisation can degrade quality more than weight quantisation.
However, I wish their testing were more detailed. Firstly, some model families are more sensitive to KV quantisation than others (only Kimi K2.6 was tested). Secondly, the evaluation suite they use to claim that FP8 KV quantisation is indistinguishable is noticeably lacking coding benchmarks; in long-running tasks, minor tool call errors compound over time.
- syntaxing - 3692 sekunder sedan> View pricing in the Cloudflare dashboard ↗
Why… I wanted to see if it’s worth it to use cloudflare’s endpoint but I can’t even see the pricing
- brokenodo - 7116 sekunder sedanI was interested in reading this until my slop detector went off at the paragraph starting with “It's worth being precise about where the benefit comes from, because it isn't raw speed.”
I love AI, but I really hate reading it.
Nördnytt! 🤓