alt.hn

8/4/2026 at 3:08:01 PM

Benchmarking Fable, Sol, and Kimi K3 on SlopCodeBench

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/benchmarking-sol-fable-kimi-on-slop-code-bench.md

by dhorthy

8/4/2026 at 6:39:54 PM

Two things

1. If we're using native harnesses, I'd have preferred you use kimi code, not opencode 2. The variation in the two kimi providers just shows how you can't trust n = 1 trials

by Bolwin