alt.hn

7/29/2026 at 8:59:08 PM

My Local LLM Scored 6/6. It Was Wrong Every Time

https://markbhall.dev/writing/my-local-llm-scored-6-of-6/

by jacquesm

7/30/2026 at 3:49:06 AM

Jfc Claude's writing can be really intolerable sometimes. Trying to parse wtf this post is saying - I guess the author is trying to make a harness for local models, and tried to have codex/Claude code autonomously make the harness better, used some sort of vibe coded test suite for that goal that was broken, switched to a different benchmark later, and the conclusion is that occasionally you can use an AI to over fit a harness to a benchmark?

by vikramkr