Logge's model notes

Beginner

My current main driver and a dated quality/cost comparison using Artificial Analysis data.

Last updated: Sep 13, 2026

My main driver

GPT-6 Astra

GPT-6 Astra is currently my main driver.

That is my personal choice. The figures below separately show how selected configurations perform in Artificial Analysis benchmarks.

Three more models on the cost frontier

A selection from the Artificial Analysis Intelligence Index v4.3. Each result belongs to a specific model version and reasoning effort.

Small budget

GPT-5.6 Luna (max)

37.50 index points · $0.18 per AA task

The cheapest entry in this selection. Its lower index score is a trade-off that needs to make sense for your task.

Model on Artificial Analysis

More index points per task

GLM-5.3-Flash

41.91 index points · $0.25 per AA task

Scores above Luna max at about $0.25 per task here. A candidate to test on your own work.

Model on Artificial Analysis

The highest score in this selection

Claude Fable 5.1 (max with fallback)

53.37 index points · $7.63 per AA task

Scores slightly above Astra max at more than twice the cost per task. The small score gap does not establish an advantage for your work.

Model on Artificial Analysis

Quality versus cost

13 selected configurations. A model is on the Pareto frontier when no other configuration in this selection scores at least as high and costs at most as much, with a strict improvement in at least one metric.

Artificial Analysis Intelligence Index v4.3
Intelligence Index and cost for 13 model configurationsHigher means more index points; further left means lower cost. Diamonds mark the frontier of this selection, circles the other models. The ring follows the selection below. All values are also available in the expandable table.35404550550.100.250.5012510Astra maxLuna maxGLM FlashFable max

Cost per AA Index task (USD, logarithmic)

Frontier in this selection Outperformed by another configuration Selected

The cost axis is logarithmic: equal distances represent equal price ratios. The index axis starts at 34. All points use unrounded source values.

Artificial Analysis Intelligence Index v4.3
52.81
Cost per AA Index task
$3.26

On the frontier of this selection

No other configuration shown improves either metric without making the other worse.

All 13 configurations and sources
Displayed values are rounded to two decimals. The frontier is calculated before rounding. Fallback is part of the stated Fable configuration.
Model / configurationIndexUSD / AA taskFrontier
GPT-6 Astra (low)45.99$0.82Yes
GPT-6 Astra (medium)49.67$1.54Yes
GPT-6 Astra (high)51.05$1.72Yes
GPT-6 Astra (xhigh)52.51$2.31Yes
GPT-6 Astra (max)52.81$3.26Yes
GPT-5.6 Luna (max)37.50$0.18Yes
GLM-5.3-Flash41.91$0.25Yes
Claude Fable 5.1 (xhigh with fallback)53.18$5.98Yes
Claude Fable 5.1 (max with fallback)53.37$7.63Yes
GPT-5.6 Sol (high)42.50$0.81Yes
GPT-5.6 Sol (max)47.06$1.99No
DeepSeek V4.1 Flash (max)39.55$0.27No
Muse Spark 1.3 (max)48.17$1.60No

What this comparison tells you

The index combines several benchmarks. Cost per task reflects their token usage and weighting, including reasoning and caching. It is not a quote for your prompts. The selection includes five Astra efforts, two Fable efforts, two Sol efforts, plus Luna, GLM, DeepSeek and Muse.

This frontier applies to this selection and these two metrics. Speed, local use, privacy and performance on your specific task are not included. Small score gaps may be within measurement uncertainty. Benchmark results are not personal experience reports.

AA also places all five Astra efforts on its broader cost/index frontier. AA analysis dated 9 September 2026

Data retrieved on . A fixed snapshot. No automatic live updates.

Archive: personal ranking from 22 August 2026

The earlier ranking remains here: Fable 5 in S+, Sol in A, and the separate Google row. It describes those earlier versions and my assessment at the time. My current main driver is shown above.

S+The favorite at the time

Fable 5

AExcellent — top picks for most tasks

gpt-5.6-sol

BGood — solid performers with some trade-offs

kimi k3

gpt-5.6-luna

deepseek v4 flash

CMediocre — notable flaws hold them back

Grok 4.6

Muse Spark 1.2

DDisappointing — overpromised, underdelivered

Opus 5

Composer 2.5

glm-5.3

gpt-5.6-terra

Sonnet 5

FBottom of the barrel — just don't

deepseek v4 pro

GoogleGoogle models, kept in their own row in Logge’s ranking

Gemini 3.7 Flash

Gemini 3.1 Pro