🎸 'Cause I Wanna Be Anarchy... Hey everybody, I'm going to be a panelist at Archon 49 next month. If you're planning on attending, here's where you can find me: It's Just a Step to the Left, and Then a Punch to the Right... 2 Oct 2026, Friday 14:00 - 15:00, Salon 1 (Gateway Center) Writing martial arts combat scenes for readers whose knowledge may be limited to watching MMA. A Conundrum of Crowns (Moderator) 2 Oct 2026, Friday 15:00 - 16:00, Marquette B (Gateway Center) Discuss why so many fantasy/sf governments seem to be kingdoms or autocracies. Why do authoritarians so often rule -when in either genre you could plausibly set up literally ANY kind of government? Losing Main Character Energy? 2 Oct 2026, Friday 20:00 - 21:00, Marquette A (Gateway Center) Have you ever grown bored of your main character? What can you do to make them more interesting, or is it better to kill them off? CONCERT: Uncle Fluffy's Post-Apocalyptic Sing-Along (Performer) 3 Oct 2026, ...
⚙️ "You are technically correct, the best kind of correct..." Today I ran across an article from tech blogger Dan Luu about "The Benchmark-pocalypse." Luu had experimented with an LLM-generated regex engine and tested it against Rebar, a comprehensive benchmark suite. The engine out-performed the top scorer by 40%! Only, it didn't actually do that. When he applied the same regex engine to a different benchmark suite, he found that it was on average 4x slower , and that was on the benchmarks it was actually able to complete. The LLM had gamed the numbers because it knew what benchmarking suite it was going to be compared to. In fact, Luu had gone so far as to instruct the LLM not to overfit its data to that suite... but it did it anyway. LLMs are statistical inference engines with an optimization loop. Whether you intend to or not, they will tailor their output to maximize their optimization, which means they will game the system if only because gaming metric...