AI-NATIVE · June 14, 2026
The fast model just got smart
For two years you made a trade every time you picked a model: fast and cheap, or smart and slow. Gemini 3.5 Flash just broke it. The 'Flash' tier — the cheap, quick one — now scores 55 on the Artificial Analysis Intelligence Index, ahead of Grok 4.3 and Claude Sonnet 4.6, while running over 280 tokens a second. The fast model is no longer the dumb model. That should make you re-open a decision most teams quietly froze a year ago: which model is your default, and is it still the right one? Here's how to think about it — including the catch.
For two years, picking an AI model meant making a trade. Fast and cheap, or smart and slow. You sent the hard reasoning to the big expensive model and let it grind; you sent the easy, high-volume stuff to the small quick one and accepted it would be a little dumber. Speed cost you intelligence. Everybody built around that assumption.
That assumption just took a serious hit. Google's Gemini 3.5 Flash — a Flash model, the tier that's supposed to be the cheap fast one — now scores 55 on the Artificial Analysis Intelligence Index, ahead of Grok 4.3 at 53 and Claude Sonnet 4.6 at 52, while running at over 280 output tokens per second — roughly 70% faster than the version before it. The fast model is no longer the dumb model. Let me explain why that's worth your attention, and where the catch is.
The trade you built around is weaker than it was
The whole reason you had a "smart model" and a "fast model" was that you couldn't have both. Intelligence lived at one end of the dial and speed at the other, and your architecture was really a series of bets about where on that dial each task belonged.
When a fast model posts top-tier intelligence scores, that dial stops being a straight line. You can now get near-frontier answers and frontier speed from the same call. That doesn't mean the biggest models are pointless — they still lead on the genuinely hardest reasoning. It means the gap narrowed enough that a lot of work you were routing to a slow, pricey model out of habit could now go to something three times faster without getting noticeably worse.
Your "default model" is probably a habit, not a decision
Here's the part that actually costs teams money. Most of us picked a default model sometime in 2024 or early 2025, wired it in, and never looked again. The leaderboard, meanwhile, reshuffles roughly every month. Your default is a snapshot of who was best the week you chose, frozen into your codebase.
That's an expensive thing to leave on autopilot, because the model market moves faster than almost any decision in your stack. The model that was clearly best a year ago may now be slower, dumber, and pricier than a tier you dismissed as the "cheap" option. The only way to know is to look again — and almost nobody does.
The catch: "fast" no longer means "cheapest"
Now the honest asterisk, because the headline hides it. This new wave of fast models also got meaningfully more expensive — Google followed Anthropic and OpenAI in raising prices on the newer, better models. So "Flash caught up on intelligence" does not automatically mean "Flash is the cheap choice it used to be." The tiers are scrambling: a fast model can be smart and not-cheap; an older model can be cheap and not-smart.
Which is exactly why you can't pick on reputation anymore. "Flash = cheap, Opus = smart" was a clean mental model and it's now wrong in both directions. The three things you actually care about — quality, latency, and cost — no longer move together, so you have to look at all three, for your task, with real numbers.
What to do
Re-benchmark on your own workload. Not on the leaderboard — on your actual prompts, your actual quality bar, your actual volume. Take the task you've been sending to your default model and run it against two or three current options, including a fast tier you'd previously have dismissed. Measure quality, measure latency, measure cost per call. Then decide, knowing the answer has a shelf life of maybe a quarter.
The bottom line
The speed-versus-intelligence trade that organized everyone's model choices just got a lot blurrier: the fast tier is posting top-tier scores, and the price tiers are scrambling along with it. The frontier moved. Your defaults didn't.
The model you picked a year ago is a year-old decision sitting in a market that turns over monthly — and "fast," "smart," and "cheap" are no longer the same axis. Re-open the choice, measure on your own task, and get used to doing it on a schedule. The trade you built around isn't gone, but it's no longer the one you memorized.
Comments
No comments yet
Sign in to join the conversation.
Be the first to share a thought.