Best performance on Podnami
LLM Rankings
Models from Settings → AI Selection, scored for live co-hosts: speed (TTFT), reliability, conversation quality (CQ), and credit value. High speed alone does not mean high CQ.
Last completed run: 2026-09-28T01:12:06.44881+00:00 · 22 models
Click a column header to sort.
Ranked models
| # | Model | Stars | Overall | Speed | Reliability |
CQ
?
CQ — Conversation Quality Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model. |
Value | TTFT p50 | Tier | Badges |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 – |
Gemini 2.5 Lite
Google · google/gemini-2.5-flash-lite
🔒 12 runs at #1
|
★★★★★ | 97.8 | 95 | 100 | 99 | 97 | 0.42s | Budget | Best value |
| 2 ▲1 |
Ministral 14b-2512
Ministral 14b-2512 · mistralai/ministral-14b-2512
|
★★★★★ | 96.0 | 100 | 100 | 91 | 97 | 0.35s | Budget | |
| 3 ▲1 |
Google Gemini Flash Lite 3.5
Google · google/gemini-3.5-flash-lite
|
★★★★★ | 95.7 | 100 | 100 | 99 | 75 | 0.47s | Standard | |
| 4 ▲1 |
Mistral Nemo
Mistral · mistralai/mistral-nemo
▲ best #1
|
★★★★★ | 95.7 | 95 | 100 | 94 | 96 | 0.43s | Budget | Not recommended for live co-hosts |
| 5 ▼3 |
Inclusion Ling 3.0 Flash
Inclusion · inclusionai/ling-3.0-flash
▲ best #1
|
★★★★★ | 95.6 | 95 | 95 | 96 | 95 | 0.87s | Budget | |
| 6 – |
Amazon Nova Lite V1
Amazon · amazon/nova-lite-v1
🔒 2 runs at #6
▲ best #5
|
★★★★★ | 94.9 | 100 | 100 | 89 | 96 | 0.51s | Budget | |
| 7 ▲2 |
GPT-3.5 Turbo
OpenAI · openai/gpt-3.5-turbo
▲ best #5
|
★★★★★ | 92.5 | 95 | 100 | 95 | 72 | 0.55s | Standard | |
| 8 ▲2 |
Gemma 4 31B IT
Google · google/gemma-4-31b-it
▲ best #1
|
★★★★★ | 91.4 | 80 | 100 | 95 | 89 | 0.32s | Budget | |
| 9 ▲3 |
GPT-4 Turbo
OpenAI · openai/gpt-4-turbo
|
★★★★★ | 91.3 | 95 | 100 | 99 | 54 | 1.16s | Premium | |
| 10 ▼3 |
GPT-4o Mini
OpenAI · openai/gpt-4o-mini
▲ best #4
|
★★★★★ | 89.7 | 80 | 100 | 99 | 68 | 0.56s | Standard | |
| 11 ▼3 |
Claude Haiku 4.5
Anthropic · anthropic/claude-haiku-4.5
▲ best #3
|
★★★★★ | 89.7 | 80 | 100 | 99 | 68 | 0.89s | Standard | |
| 12 ▲2 |
DeepSeek Chat
DeepSeek · deepseek/deepseek-chat
▲ best #1
|
★★★★★ | 88.3 | 65 | 100 | 99 | 84 | 1.56s | Budget | |
| 13 – |
DeepSeek V4 Flash
DeepSeek · deepseek/deepseek-v4-flash
🔒 2 runs at #13
▲ best #9
|
★★★★☆ | 81.3 | 50 | 95 | 96 | 75 | 2.80s | Budget | |
| 14 ▼3 |
DeepseekV32
DeepSeek · deepseek/deepseek-v3.2
▲ best #2
|
★★★★☆ | 79.1 | 50 | 95 | 91 | 73 | 1.50s | Budget | |
| 15 ▲1 |
Meta Llamma Scout
Meta Llamma Scout · meta-llama/llama-4-scout
▲ best #2
|
★★★☆☆ | 56.7 | 95 | 32 | 42 | 64 | 0.47s | Budget | Unstable |
| 16 ▲2 |
Qwen 3.7 Flash
OpenRouter · qwen/qwen3.7-flash
|
★☆☆☆☆ | 38.3 | 80 | 9 | 24 | 46 | 2.32s | Budget | Unstable |
| 17 ▼2 |
Step 3.5 Flash
StepFun · stepfun/step-3.5-flash
▲ best #15
|
★☆☆☆☆ | 10.8 | 0 | 0 | 24 | 8 | — | Budget | Unstable |
| 18 ▼1 |
Inclusion Ling 2.6 Flash
Ling · inclusionai/ling-2.6-flash
▲ best #17
|
★☆☆☆☆ | 10.8 | 0 | 0 | 24 | 8 | — | Budget | No endpoint |
| 19 – |
Claude 3 Haiku
Anthropic · anthropic/claude-3-haiku
🔒 3 runs at #19
▲ best #2
|
★☆☆☆☆ | 10.4 | 0 | 0 | 24 | 6 | — | Standard | No endpoint |
| 20 – |
Incllusion Ling 2.6
Inclusion · inclusionai/ling-2.6-1t
🔒 30 runs at #20
|
★☆☆☆☆ | 9.1 | 0 | 0 | 20 | 7 | — | Budget | No endpoint |
Free options (comparison)
| # | Model | Stars | Overall | Speed | Reliability |
CQ
?
CQ — Conversation Quality Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model. |
Value | TTFT p50 | Tier | Badges |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 ▲1 |
FREE LLM
OpenRouter · free-llm-cycler
|
★★★★☆ | 76.0 | 65 | 86 | 79 | 74 | 1.28s | Free | Free option Not recommended for live co-hosts |
| 2 ▼1 |
Free LLM Router
Open Router · openrouter/free
▲ best #1
|
★★★☆☆ | 69.2 | 65 | 86 | 64 | 69 | 1.36s | Free | Free option Not recommended for live co-hosts |
▲/▼ = rank change since the previous run · – = no change · NEW = first appearance
Probes are scripted OpenRouter checks (not live mixer sessions). Pass/fail + timings only — no reply transcripts stored. Scores reflect a sampled run (a few attempts per model at a non-zero temperature), so small differences between runs are normal — treat close scores as ties, not regressions.
