Best performance on Podnami
LLM Rankings
Models from Settings → AI Selection, scored for live co-hosts: speed (TTFT), reliability, conversation quality (CQ), and credit value. High speed alone does not mean high CQ.
Last completed run: 2026-09-12T01:14:25.102839+00:00 · 22 models
Click a column header to sort.
Ranked models
| # | Model | Stars | Overall | Speed | Reliability |
CQ
?
CQ — Conversation Quality Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model. |
Value | TTFT p50 | Tier | Badges |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 ▲1 |
Gemini 2.5 Lite
Google · google/gemini-2.5-flash-lite
|
★★★★★ | 97.8 | 95 | 100 | 99 | 97 | 0.38s | Budget | Best value |
| 2 ▲1 |
Inclusion Ling 3.0 Flash
Inclusion · inclusionai/ling-3.0-flash
▲ best #1
|
★★★★★ | 97.8 | 95 | 100 | 99 | 97 | 1.09s | Budget | |
| 3 ▲1 |
Ministral 14b-2512
Ministral 14b-2512 · mistralai/ministral-14b-2512
▲ best #2
|
★★★★★ | 96.2 | 95 | 100 | 95 | 96 | 0.34s | Budget | |
| 4 ▲10 |
Claude 3 Haiku
Anthropic · anthropic/claude-3-haiku
|
★★★★★ | 95.7 | 100 | 100 | 99 | 75 | 0.38s | Standard | |
| 5 – |
Claude Haiku 4.5
Anthropic · anthropic/claude-haiku-4.5
🔒 2 runs at #5
▲ best #3
|
★★★★★ | 95.7 | 100 | 100 | 99 | 75 | 0.71s | Standard | |
| 6 ▲3 |
Amazon Nova Lite V1
Amazon · amazon/nova-lite-v1
▲ best #4
|
★★★★★ | 94.9 | 100 | 100 | 89 | 96 | 0.44s | Budget | |
| 7 – |
GPT-3.5 Turbo
OpenAI · openai/gpt-3.5-turbo
🔒 2 runs at #7
▲ best #5
|
★★★★★ | 94.0 | 100 | 100 | 95 | 74 | 0.76s | Standard | |
| 8 – |
Google Gemini Flash Lite 3.5
Google · google/gemini-3.5-flash-lite
🔒 3 runs at #8
▲ best #3
|
★★★★★ | 94.0 | 100 | 100 | 95 | 74 | 0.51s | Standard | |
| 9 ▲6 |
DeepSeek V4 Flash
DeepSeek · deepseek/deepseek-v4-flash
|
★★★★★ | 93.1 | 80 | 100 | 99 | 91 | 1.27s | Budget | |
| 10 ▲2 |
DeepseekV32
DeepSeek · deepseek/deepseek-v3.2
▲ best #2
|
★★★★★ | 91.4 | 80 | 100 | 95 | 89 | 0.88s | Budget | |
| 11 ▲2 |
GPT-4 Turbo
OpenAI · openai/gpt-4-turbo
|
★★★★★ | 91.3 | 95 | 100 | 99 | 54 | 1.29s | Premium | |
| 12 ▼6 |
GPT-4o Mini
OpenAI · openai/gpt-4o-mini
▲ best #5
|
★★★★★ | 89.7 | 80 | 100 | 99 | 68 | 0.50s | Standard | |
| 13 ▼2 |
Gemma 4 31B IT
Google · google/gemma-4-31b-it
▲ best #1
|
★★★★★ | 89.7 | 80 | 100 | 91 | 88 | 0.24s | Budget | |
| 14 ▼4 |
Mistral Nemo
Mistral · mistralai/mistral-nemo
▲ best #1
|
★★★★★ | 89.2 | 80 | 100 | 90 | 88 | 1.11s | Budget | Not recommended for live co-hosts |
| 15 ▲1 |
Step 3.5 Flash
StepFun · stepfun/step-3.5-flash
|
★★★☆☆ | 68.3 | 60 | 72 | 72 | 67 | 2.87s | Budget | |
| 16 ▼15 |
DeepSeek Chat
DeepSeek · deepseek/deepseek-chat
▲ best #1
|
★★★☆☆ | 61.7 | 65 | 43 | 69 | 62 | 1.43s | Budget | Unstable |
| 17 – |
Meta Llamma Scout
Meta Llamma Scout · meta-llama/llama-4-scout
🔒 2 runs at #17
▲ best #2
|
★★☆☆☆ | 52.0 | 80 | 32 | 42 | 57 | 0.20s | Budget | Unstable |
| 18 – |
Qwen 3.7 Flash
OpenRouter · qwen/qwen3.7-flash
🔒 6 runs at #18
|
★☆☆☆☆ | 37.2 | 80 | 5 | 24 | 45 | 2.03s | Budget | Unstable |
| 19 – |
Inclusion Ling 2.6 Flash
Ling · inclusionai/ling-2.6-flash
🔒 6 runs at #19
▲ best #2
|
★☆☆☆☆ | 10.8 | 0 | 0 | 24 | 8 | — | Budget | No endpoint |
| 20 – |
Incllusion Ling 2.6
Inclusion · inclusionai/ling-2.6-1t
🔒 19 runs at #20
▲ best #2
|
★☆☆☆☆ | 9.1 | 0 | 0 | 20 | 7 | — | Budget | No endpoint |
Free options (comparison)
| # | Model | Stars | Overall | Speed | Reliability |
CQ
?
CQ — Conversation Quality Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model. |
Value | TTFT p50 | Tier | Badges |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 ▲1 |
FREE LLM
OpenRouter · free-llm-cycler
|
★★★★☆ | 76.8 | 50 | 95 | 86 | 72 | 2.06s | Free | Free option |
| 2 ▼1 |
Free LLM Router
Open Router · openrouter/free
▲ best #1
|
★★★☆☆ | 63.7 | 35 | 91 | 70 | 58 | 2.58s | Free | Free option Not recommended for live co-hosts |
▲/▼ = rank change since the previous run · – = no change · NEW = first appearance
Probes are scripted OpenRouter checks (not live mixer sessions). Pass/fail + timings only — no reply transcripts stored. Scores reflect a sampled run (a few attempts per model at a non-zero temperature), so small differences between runs are normal — treat close scores as ties, not regressions.
