Benchmarks Misbehave, Prices Move, and Databricks Buys Again

Grok 4.6 Has an Effort Problem
Oh, this is deliciously inconvenient. The headline floating around the DeepSWE board is that Grok 4.6 /medium outperforms /high effort. That’s not just a leaderboard quirk; that’s a tiny grenade rolled under the whole “more thinking equals better answers” sales pitch.
One camp will say: relax, benchmarks are weird, single results can mislead, and DeepSWE is only one arena. Fair. But the other camp has the sharper TV line: if “high effort” costs more time, more compute, or more patience, shouldn’t it actually win? If the middle setting beats the fancy setting, users are going to ask why they’re being nudged toward the expensive button with the serious-sounding label.

The stakes are simple: model makers want effort knobs to feel like premium control panels. Skeptics see them as vibes with a price tag. Today, the skeptics get the better monologue.
DeepSeek Turns the Pricing Screw — and Opens the Toolbox
DeepSeek has two fresh sparks on the table: a GitHub project called DeepSeek Harness and an API Pricing Update. That combination is catnip for the AI crowd because it hits both sides of the developer brain: “Can I test this properly?” and “What’s this going to cost me?”
The debate practically writes itself. Builders want predictable pricing and clean tooling. Rival model shops don’t want DeepSeek to become the default cheap-and-serious option. And finance teams? They’re standing behind every engineer whispering, “Please don’t surprise me with another mystery bill.”

