Best performance on Podnami
LLM Rankings
Models from Settings → AI Selection, scored for live co-hosts: speed (TTFT), reliability, conversation quality (CQ), and credit value. High speed alone does not mean high CQ.
Last completed run: 2026-08-14T01:17:09.334537+00:00 · 19 models
Click a column header to sort.
Ranked models
| # | Model | Stars | Overall | Speed | Reliability |
CQ
?
CQ — Conversation Quality Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model. |
Value | TTFT p50 | Tier | Badges |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 ▲1 |
Gemini 2.5 Lite
Google · google/gemini-2.5-flash-lite
|
★★★★★ | 99.4 | 100 | 100 | 99 | 100 | 0.45s | Budget | Best value |
| 2 ▲11 |
Inclusion Ling 2.6 Flash
Ling · inclusionai/ling-2.6-flash
|
★★★★★ | 97.8 | 95 | 100 | 99 | 97 | 0.54s | Budget | |
| 3 ▲3 |
Google Gemini Flash Lite 3.5
Google · google/gemini-3.5-flash-lite
|
★★★★★ | 95.7 | 100 | 100 | 99 | 75 | 0.42s | Standard | |
| 4 ▲11 |
Mistral Nemo
Mistral · mistralai/mistral-nemo
|
★★★★★ | 95.7 | 95 | 100 | 94 | 96 | 0.45s | Budget | Not recommended for live co-hosts |
| 5 ▲9 |
GPT-3.5 Turbo
OpenAI · openai/gpt-3.5-turbo
▲ best #4
|
★★★★★ | 94.0 | 100 | 100 | 95 | 74 | 0.31s | Standard | |
| 6 ▲2 |
Amazon Nova Lite V1
Amazon · amazon/nova-lite-v1
▲ best #1
|
★★★★★ | 93.3 | 95 | 100 | 89 | 94 | 0.47s | Budget | |
| 7 ▼6 |
DeepSeek Chat
DeepSeek · deepseek/deepseek-chat
▲ best #1
|
★★★★★ | 93.1 | 80 | 100 | 99 | 91 | 1.57s | Budget | |
| 8 ▲1 |
DeepseekV32
DeepSeek · deepseek/deepseek-v3.2
▲ best #3
|
★★★★★ | 91.4 | 80 | 100 | 95 | 89 | 1.23s | Budget | |
| 9 ▲1 |
Ministral 14b-2512
Ministral 14b-2512 · mistralai/ministral-14b-2512
▲ best #2
|
★★★★★ | 91.4 | 80 | 100 | 95 | 89 | 0.44s | Budget | |
| 10 ▼3 |
Meta Llamma Scout
Meta Llamma Scout · meta-llama/llama-4-scout
▲ best #2
|
★★★★★ | 91.4 | 80 | 100 | 95 | 89 | 0.31s | Budget | |
| 11 – |
GPT-4 Turbo
OpenAI · openai/gpt-4-turbo
🔒 2 runs at #11
|
★★★★★ | 91.3 | 95 | 100 | 99 | 54 | 1.42s | Premium | |
| 12 ▼9 |
GPT-4o Mini
OpenAI · openai/gpt-4o-mini
▲ best #3
|
★★★★★ | 89.7 | 80 | 100 | 99 | 68 | 0.51s | Standard | |
| 13 ▼9 |
Claude 3 Haiku
Anthropic · anthropic/claude-3-haiku
▲ best #4
|
★★★★★ | 89.7 | 80 | 100 | 99 | 68 | 0.34s | Standard | |
| 14 ▲2 |
Gemma 4 31B IT
Google · google/gemma-4-31b-it
▲ best #3
|
★★★★★ | 89.7 | 80 | 100 | 91 | 88 | 0.34s | Budget | |
| 15 ▼3 |
DeepSeek V4 Flash
DeepSeek · deepseek/deepseek-v4-flash
▲ best #12
|
★★★★★ | 86.1 | 65 | 95 | 96 | 82 | 1.63s | Budget | |
| 16 ▼11 |
Claude Haiku 4.5
Anthropic · anthropic/claude-haiku-4.5
▲ best #5
|
★★★★★ | 85.2 | 65 | 100 | 99 | 63 | 0.89s | Standard | |
| 17 – |
Step 3.5 Flash
StepFun · stepfun/step-3.5-flash
🔒 8 runs at #17
|
★★★☆☆ | 58.0 | 45 | 62 | 65 | 55 | 2.66s | Budget |
Free options (comparison)
| # | Model | Stars | Overall | Speed | Reliability |
CQ
?
CQ — Conversation Quality Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model. |
Value | TTFT p50 | Tier | Badges |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 – |
FREE LLM
OpenRouter · free-llm-cycler
🔒 3 runs at #1
|
★★★★☆ | 70.3 | 45 | 100 | 73 | 66 | 1.78s | Free | Free option Not recommended for live co-hosts |
| 2 – |
Free LLM Router
Open Router · openrouter/free
🔒 3 runs at #2
▲ best #1
|
★★★☆☆ | 66.8 | 45 | 90 | 70 | 63 | 2.46s | Free | Free option Not recommended for live co-hosts |
▲/▼ = rank change since the previous run · – = no change · NEW = first appearance
Probes are scripted OpenRouter checks (not live mixer sessions). Pass/fail + timings only — no reply transcripts stored. Scores reflect a sampled run (a few attempts per model at a non-zero temperature), so small differences between runs are normal — treat close scores as ties, not regressions.
