How I found out my agent's memory was corrupted
A traveler asked why my agent kept sending her hostels. It had saved a colleague's sixty-euro budget as her own preference weeks earlier and had been following it ever since. Here is how I found it, and why no single trace showed it.
Coding ability is not what separates frontier models
Five frontier models ran the same 410 coding-agent tasks, traced into Latitude. Capability tied; cost, speed, and refusals did not. What actually separates them and how to choose.
Reasoning effort only matters when the agent can't run tests
A Reddit post argued Opus 5 codes worse at higher reasoning effort. I swept all five effort levels, 155 runs. With tests available every level went 18 for 18. With tests hidden, solves went from 0 to 8 out of 13.
