Podnami

Best performance on Podnami

LLM Rankings

Models from Settings → AI Selection, scored for live co-hosts: speed (TTFT), reliability, conversation quality (CQ), and credit value. High speed alone does not mean high CQ.

Last completed run: 2026-09-28T01:12:06.44881+00:00 · 22 models

Click a column header to sort.

Ranked models

# Model Stars Overall Speed Reliability CQ ? CQ — Conversation Quality
Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model.
Value TTFT p50 Tier Badges
1 –
Gemini 2.5 Lite
Google · google/gemini-2.5-flash-lite
🔒 12 runs at #1
★★★★★ 97.8 95 100 99 97 0.42s Budget Best value
2 ▲1
Ministral 14b-2512
Ministral 14b-2512 · mistralai/ministral-14b-2512
★★★★★ 96.0 100 100 91 97 0.35s Budget
3 ▲1
Google Gemini Flash Lite 3.5
Google · google/gemini-3.5-flash-lite
★★★★★ 95.7 100 100 99 75 0.47s Standard
4 ▲1
Mistral Nemo
Mistral · mistralai/mistral-nemo
▲ best #1
★★★★★ 95.7 95 100 94 96 0.43s Budget Not recommended for live co-hosts
5 ▼3
Inclusion Ling 3.0 Flash
Inclusion · inclusionai/ling-3.0-flash
▲ best #1
★★★★★ 95.6 95 95 96 95 0.87s Budget
6 –
Amazon Nova Lite V1
Amazon · amazon/nova-lite-v1
🔒 2 runs at #6 ▲ best #5
★★★★★ 94.9 100 100 89 96 0.51s Budget
7 ▲2
GPT-3.5 Turbo
OpenAI · openai/gpt-3.5-turbo
▲ best #5
★★★★★ 92.5 95 100 95 72 0.55s Standard
8 ▲2
Gemma 4 31B IT
Google · google/gemma-4-31b-it
▲ best #1
★★★★★ 91.4 80 100 95 89 0.32s Budget
9 ▲3
GPT-4 Turbo
OpenAI · openai/gpt-4-turbo
★★★★★ 91.3 95 100 99 54 1.16s Premium
10 ▼3
GPT-4o Mini
OpenAI · openai/gpt-4o-mini
▲ best #4
★★★★★ 89.7 80 100 99 68 0.56s Standard
11 ▼3
Claude Haiku 4.5
Anthropic · anthropic/claude-haiku-4.5
▲ best #3
★★★★★ 89.7 80 100 99 68 0.89s Standard
12 ▲2
DeepSeek Chat
DeepSeek · deepseek/deepseek-chat
▲ best #1
★★★★★ 88.3 65 100 99 84 1.56s Budget
13 –
DeepSeek V4 Flash
DeepSeek · deepseek/deepseek-v4-flash
🔒 2 runs at #13 ▲ best #9
★★★★☆ 81.3 50 95 96 75 2.80s Budget
14 ▼3
DeepseekV32
DeepSeek · deepseek/deepseek-v3.2
▲ best #2
★★★★☆ 79.1 50 95 91 73 1.50s Budget
15 ▲1
Meta Llamma Scout
Meta Llamma Scout · meta-llama/llama-4-scout
▲ best #2
★★★☆☆ 56.7 95 32 42 64 0.47s Budget Unstable
16 ▲2
Qwen 3.7 Flash
OpenRouter · qwen/qwen3.7-flash
★☆☆☆☆ 38.3 80 9 24 46 2.32s Budget Unstable
17 ▼2
Step 3.5 Flash
StepFun · stepfun/step-3.5-flash
▲ best #15
★☆☆☆☆ 10.8 0 0 24 8 — Budget Unstable
18 ▼1
Inclusion Ling 2.6 Flash
Ling · inclusionai/ling-2.6-flash
▲ best #17
★☆☆☆☆ 10.8 0 0 24 8 — Budget No endpoint
19 –
Claude 3 Haiku
Anthropic · anthropic/claude-3-haiku
🔒 3 runs at #19 ▲ best #2
★☆☆☆☆ 10.4 0 0 24 6 — Standard No endpoint
20 –
Incllusion Ling 2.6
Inclusion · inclusionai/ling-2.6-1t
🔒 30 runs at #20
★☆☆☆☆ 9.1 0 0 20 7 — Budget No endpoint

Free options (comparison)

# Model Stars Overall Speed Reliability CQ ? CQ — Conversation Quality
Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model.
Value TTFT p50 Tier Badges
1 ▲1
FREE LLM
OpenRouter · free-llm-cycler
★★★★☆ 76.0 65 86 79 74 1.28s Free Free option Not recommended for live co-hosts
2 ▼1
Free LLM Router
Open Router · openrouter/free
▲ best #1
★★★☆☆ 69.2 65 86 64 69 1.36s Free Free option Not recommended for live co-hosts

▲/▼ = rank change since the previous run · – = no change · NEW = first appearance

Probes are scripted OpenRouter checks (not live mixer sessions). Pass/fail + timings only — no reply transcripts stored. Scores reflect a sampled run (a few attempts per model at a non-zero temperature), so small differences between runs are normal — treat close scores as ties, not regressions.