How real is the 4x price difference?
Completely real, and at this tier it compounds fast. GPT-5 Mini: $0.25/Mtok in, $2/Mtok out. Haiku 4.5: $1 in, $5 out. On 1M input and 200K output per day, Mini costs about $19.50/month against Haiku's $60. If your cheap tier handles ten times that volume, you're comparing $195 against $600 monthly for the same call pattern.
Both are reasoning-capable, both do tool calling, both take text and images. Haiku adds PDF input; Mini counters with a 400K context window, double Haiku's 200K.
What does Haiku's premium buy?
Three things worth naming precisely. First, recency: Haiku 4.5 shipped October 2025, GPT-5 Mini in August 2025, and at the small-model tier generational gains land noticeably.
Second, code quality. For a budget model, Haiku's code output is unusually strong, closer to what you'd expect a tier up. If your cheap calls generate snippets, tests, or config, that shows.
Third, and most important for agent builders: behavior under multi-step tool use. Haiku degrades more gracefully when a task turns out to be harder than its router thought. Mini is excellent at genuinely mechanical work but falls off faster when a "simple" call quietly requires judgment.
Which handles longer prompts?
GPT-5 Mini, on paper: 400K tokens of context against Haiku's 200K. If your cheap tier stuffs large documents or long histories into prompts, check your P95 prompt length before choosing; being forced to chunk at 200K when your workload occasionally needs 300K is an architectural annoyance no price difference fixes.
How should you choose for an agent pipeline?
Split your cheap-tier calls into two buckets and be honest about which is which.
Mechanical bucket, where the task is fully specified and the schema is tight: moderation, tagging, dedup, field extraction, formatting. Route these to GPT-5 Mini. At these tasks the models are interchangeable and Mini is four times cheaper.
Judgment bucket, where calls occasionally need real thinking: triage that considers context, extraction from messy input, substeps that chain tools. Route these to Haiku 4.5. The premium disappears the first time you count what retries and bad outputs cost downstream.
Teams that route this way typically end up with most calls on Mini and a minority on Haiku, paying budget prices for bulk work while keeping quality where it matters. The mistake is forcing one model to cover both buckets; either you overpay for the mechanical work or you under-serve the judgment work.