⚙️ "You are technically correct, the best kind of correct..." Today I ran across an article from tech blogger Dan Luu about "The Benchmark-pocalypse." Luu had experimented with an LLM-generated regex engine and tested it against Rebar, a comprehensive benchmark suite. The engine out-performed the top scorer by 40%! Only, it didn't actually do that. When he applied the same regex engine to a different benchmark suite, he found that it was on average 4x slower , and that was on the benchmarks it was actually able to complete. The LLM had gamed the numbers because it knew what benchmarking suite it was going to be compared to. In fact, Luu had gone so far as to instruct the LLM not to overfit its data to that suite... but it did it anyway. LLMs are statistical inference engines with an optimization loop. Whether you intend to or not, they will tailor their output to maximize their optimization, which means they will game the system if only because gaming metric...
⚡ You're burnin' up the quarter mile... It's that time of year once again where Kurt spends an inordinate amount of time at a conference playing board games. There were some new titles, some returning favorites, and a whoooooooole lot of people. This year, instead of trying to play everything that looked interesting to me, I went a little deeper on fewer titles. I would have loved to try out some of the heavier games, but while I had the desire, I had not the mental fortitude to learn the rules ahead of time. Alas. Better luck next year. The themes for this year seemed to be trains, trains, tableau-builders, and more trains. Seriously, there were I think six different train games here, including Ticket to Ride and two that just started with the word Railroad . Anyway, without further ado, here are the new-to-me titles that I played at Geekway To The West 2026! Lightning Train This the was the first game I played, it was easily my favorite of the con, and I won a copy in th...