Skip to content
ZAIQZAIQ

Zaiq AI Benchmark · September 2026

Which AI model is best? The Zaiq AI Benchmark

Independent benchmark results for the AI models that matter, turned into plain answers: which is the most capable, which makes up facts least, what each costs in rand, and which to use for what. Updated every month.

Chad AlexanderCo-founder and AI engineerUpdated 28 September 2026

The short answer

Most capable: GPT-6 Astra (166.6) and Claude Fable 5.1 (165) are effectively tied at the top: their uncertainty ranges overlap.

Best value: Gemini 3.8 Flash scores within 10 points of the leader and costs about R31 to answer 1,000 customer messages, against R408 for the top model.

Cheapest good model: GPT-5.6 Luna, the model behind free ChatGPT, at about R9,14 per 1,000 messages.

Fewest made-up facts: GPT-6 Astra, with 75.6% of hard factual questions right without search. Best model you can run yourself: Kimi K3.

166.6
Top capability score, GPT-6 Astra
45×
Price gap between the cheapest and the most expensive model here
268
Models tracked by Epoch AI, the source of the scores

The leaderboard

Ranked by the Epoch Capabilities Index. Prices are what the model costs to use through an API, converted at US$1 = R16.32 (28 September 2026). Apps like ChatGPT and Claude charge a monthly plan instead.

#ModelCapabilityFacts rightScienceRand per million tokens (in / out)1,000 messagesReads at once
1GPT-6 AstraOpenAIChatGPT Pro (from R1,839 a month) uses it for Pro reasoning166.675.6%95.8%R163,20 / R816,00R408about 1 550 pages
2Claude Fable 5.1Anthropic16570.8%not yet testedR163,20 / R816,00R408about 1 500 pages
3Claude Opus 5Anthropic162.6759.9%93.9%R81,60 / R408,00R204about 1 500 pages
4GPT-5.6 SolOpenAI161.9969.7%93.5%R32,64 / R163,20R82about 1 550 pages
5Kimi K3Moonshot AI · open weights157.6850.6%93.1%R48,96 / R244,80R122about 1 550 pages
6Gemini 3.8 FlashGoogle157.1369.7%95.4%R12,24 / R61,20R31about 1 550 pages
7Muse Spark 1.3Meta156.89not yet testednot yet testedR20,40 / R69,36R41about 1 550 pages
8Grok 4.6xAI156.4849.3%94%R32,64 / R97,92R62about 750 pages
9Claude Sonnet 5Anthropic156.3433.7%90.5%R32,64 / R163,20R82about 1 500 pages
10GPT-5.6 LunaOpenAIChatGPT Free runs it for unlimited everyday chats156.3241%91.6%R3,26 / R19,58R9,14about 1 550 pages
11GLM-5.3Z.ai155.5641%90.9%R22,85 / R71,81R44about 1 950 pages
12DeepSeek V4 Pro 0813DeepSeek · open weights155.3952.9%91.7%R6,52 / R68,54R27about 1 550 pages
13Qwen3.8 Max (0902)Alibaba155.2847.3%92.3%R32,64 / R97,92R62about 1 500 pages
14DeepSeek V4.1 FlashDeepSeek · open weights155.01not yet testednot yet testedR4,90 / R19,58R10,77about 1 550 pages
15Gemini 3.1 ProGoogle154.92not yet testednot yet testedR32,64 / R195,84R91about 1 550 pages
16Claude Haiku 4.5Anthropic142.4213.2%71.2%R16,32 / R81,60R41about 300 pages
17Mistral Medium 3.5Mistral AI141.42not yet testednot yet testedR24,48 / R122,40R61about 400 pages

Download the full table: zaiq-ai-benchmark-2026-09.csv. Scores from Epoch AI (CC BY 4.0). Prices from OpenRouter, checked 28 September 2026.

Newer models not yet scored

These are on sale but Epoch AI has not scored them yet, so they are not ranked above. Other trackers disagree at the very top: Artificial Analysis puts Claude Opus 5.5 first on its own index, while Epoch puts GPT-6 Astra first. We add each model to the ranking as soon as Epoch publishes its score.

ModelWhat to knowRand per million tokens (in / out)1,000 messages
Claude Opus 5.5AnthropicAnthropic's newest Opus. Artificial Analysis puts it first on its own Intelligence Index: 58, ahead of Claude Fable 5.1 and GPT-6 Astra on 53 (checked 28 September 2026).R65,28 / R326,40R163
GPT-6 SolOpenAIPart of the GPT-6 family behind ChatGPT Plus.R32,64 / R163,20R82
GPT-6 LunaOpenAIThe low-cost GPT-6 model.R1,63 / R8,16R4,08
Grok 4.7xAIThe newest Grok in xAI's API.R26,11 / R78,34R50

What the numbers mean

  • Capability is the Epoch Capabilities Index. Epoch AI combines results from dozens of benchmarks, such as maths, coding, science and reasoning tests, into one scale. Higher is better. A gap of a point or two is within the margin of error; ten points is a clear difference.
  • Facts right is SimpleQA Verified: short factual questions with one correct answer, asked without web search. It shows how often a model states a fact correctly instead of guessing. Lower scores mean more made-up answers.
  • Science is GPQA Diamond: graduate-level biology, physics and chemistry questions written by experts. The top models now score above 90%, so it no longer separates them much.
  • Rand per million tokensis what the model costs through its API. A token is roughly three quarters of a word, so a million tokens is about 1,500 pages. “In” is what you send; “out” is what the model writes back.
  • 1,000 messages is our worked example of a customer-service assistant: each message sends about 1 000 tokens (instructions, context and the question) and gets about 300back. Reasoning models can “think” before answering and bill that thinking as output, so hard tasks can cost several times more.
  • Reads at once is the context window: how much text the model can take in at one time, in pages of about 500 words.

Which model for which job

  • The hardest work (complex analysis, long reports, difficult code): GPT-6 Astra or Claude Fable 5.1. They cost the most, so use them where quality pays for itself.
  • A customer assistant or document Q&A: Gemini 3.8 Flash or GPT-5.6 Luna. They are capable enough for most business questions and cost a fraction of the top models, which matters when thousands of messages come in.
  • Keeping data on your own servers: an open-weights model such as Kimi K3, or DeepSeek's V4 models, can run on infrastructure you control, which helps with POPIA's rules on sending personal information abroad. It takes engineering to run well.
  • Everyday use for free: free ChatGPT runs GPT-5.6 Luna. See ChatGPT free vs paid in South Africa for what each plan costs in rand.

Picking the model is the easy part. The value comes from connecting it to your own data and processes, with a person checking anything that touches money or customers. That is what ZAIQ builds.

How the Zaiq AI Benchmark is built

We do not run the tests ourselves. The capability, accuracy and science scores come from Epoch AI, an independent research group that publishes its benchmark results openly under a Creative Commons Attribution licence (Epoch AI, ‘Capabilities & benchmarking’, retrieved 28 September 2026). Prices and context windows come from OpenRouter's public model list, converted to rand at US$1 = R16.32. For models Epoch has not scored yet, we note where Artificial Analysis places them on its own Intelligence Index.

Our part is choosing the models South Africans actually use, matching each one across both sources, working out what it costs in rand, and explaining it in plain English. Where Epoch has not yet tested a model on a benchmark, the table says so rather than guessing. The table is rebuilt every month from the same sources, and the previous editions stay available as downloads.

Questions about the benchmark

Which AI model is the best right now?

On the Epoch Capabilities Index, GPT-6 Astra scores 166.6 and Claude Fable 5.1 165. Their uncertainty ranges overlap, so treat them as tied at the top. Both are expensive to use through an API: about R408 to answer 1,000 typical customer messages.

What is the best value AI model?

Gemini 3.8 Flash. It scores 157.13, within 10 points of the leader, and costs about R31 per 1,000 customer messages at 28 September 2026 prices. For most business assistants that answer questions from your own documents, a model in this bracket is enough.

Which AI makes up the fewest facts?

GPT-6 Astra, which answered 75.6% of the SimpleQA Verified questions correctly without web search, the highest of the models we track. Every model still gets facts wrong, so connect it to your own documents or to search for anything that matters.

Is the free ChatGPT model any good?

Free ChatGPT runs GPT-5.6 Luna, which scores 156.32 on the capability index, about 94% of the leader's score. It is also the cheapest model here through the API, at about R9,14 per 1,000 messages. For everyday writing and questions it is good enough; paid plans unlock the top models.

How much does it cost to run AI in a South African business?

For a customer-service assistant, the model itself is cheap: from about R9,14 to R408 per 1,000 messages on the models here, at US$1 = R16.32. The larger cost is building it properly into your systems. Zaiq quotes that as one fixed price in rand.

Want the right model working in your business?

ZAIQ picks the model that fits the job and the budget, connects it to your own data and systems, and quotes one fixed price in rand before any work starts.

Tell us the job→