Podnami

Best performance on Podnami

LLM Rankings

Models from Settings → AI Selection, scored for live co-hosts: speed (TTFT), reliability, conversation quality (CQ), and credit value. High speed alone does not mean high CQ.

Last completed run: 2026-09-12T01:14:25.102839+00:00 · 22 models

Click a column header to sort.

Ranked models

# Model Stars Overall Speed Reliability CQ ? CQ — Conversation Quality
Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model.
Value TTFT p50 Tier Badges
1 ▲1
Gemini 2.5 Lite
Google · google/gemini-2.5-flash-lite
★★★★★ 97.8 95 100 99 97 0.38s Budget Best value
2 ▲1
Inclusion Ling 3.0 Flash
Inclusion · inclusionai/ling-3.0-flash
▲ best #1
★★★★★ 97.8 95 100 99 97 1.09s Budget
3 ▲1
Ministral 14b-2512
Ministral 14b-2512 · mistralai/ministral-14b-2512
▲ best #2
★★★★★ 96.2 95 100 95 96 0.34s Budget
4 ▲10
Claude 3 Haiku
Anthropic · anthropic/claude-3-haiku
★★★★★ 95.7 100 100 99 75 0.38s Standard
5
Claude Haiku 4.5
Anthropic · anthropic/claude-haiku-4.5
🔒 2 runs at #5 ▲ best #3
★★★★★ 95.7 100 100 99 75 0.71s Standard
6 ▲3
Amazon Nova Lite V1
Amazon · amazon/nova-lite-v1
▲ best #4
★★★★★ 94.9 100 100 89 96 0.44s Budget
7
GPT-3.5 Turbo
OpenAI · openai/gpt-3.5-turbo
🔒 2 runs at #7 ▲ best #5
★★★★★ 94.0 100 100 95 74 0.76s Standard
8
Google Gemini Flash Lite 3.5
Google · google/gemini-3.5-flash-lite
🔒 3 runs at #8 ▲ best #3
★★★★★ 94.0 100 100 95 74 0.51s Standard
9 ▲6
DeepSeek V4 Flash
DeepSeek · deepseek/deepseek-v4-flash
★★★★★ 93.1 80 100 99 91 1.27s Budget
10 ▲2
DeepseekV32
DeepSeek · deepseek/deepseek-v3.2
▲ best #2
★★★★★ 91.4 80 100 95 89 0.88s Budget
11 ▲2
GPT-4 Turbo
OpenAI · openai/gpt-4-turbo
★★★★★ 91.3 95 100 99 54 1.29s Premium
12 ▼6
GPT-4o Mini
OpenAI · openai/gpt-4o-mini
▲ best #5
★★★★★ 89.7 80 100 99 68 0.50s Standard
13 ▼2
Gemma 4 31B IT
Google · google/gemma-4-31b-it
▲ best #1
★★★★★ 89.7 80 100 91 88 0.24s Budget
14 ▼4
Mistral Nemo
Mistral · mistralai/mistral-nemo
▲ best #1
★★★★★ 89.2 80 100 90 88 1.11s Budget Not recommended for live co-hosts
15 ▲1
Step 3.5 Flash
StepFun · stepfun/step-3.5-flash
★★★☆☆ 68.3 60 72 72 67 2.87s Budget
16 ▼15
DeepSeek Chat
DeepSeek · deepseek/deepseek-chat
▲ best #1
★★★☆☆ 61.7 65 43 69 62 1.43s Budget Unstable
17
Meta Llamma Scout
Meta Llamma Scout · meta-llama/llama-4-scout
🔒 2 runs at #17 ▲ best #2
★★☆☆☆ 52.0 80 32 42 57 0.20s Budget Unstable
18
Qwen 3.7 Flash
OpenRouter · qwen/qwen3.7-flash
🔒 6 runs at #18
★☆☆☆☆ 37.2 80 5 24 45 2.03s Budget Unstable
19
Inclusion Ling 2.6 Flash
Ling · inclusionai/ling-2.6-flash
🔒 6 runs at #19 ▲ best #2
★☆☆☆☆ 10.8 0 0 24 8 Budget No endpoint
20
Incllusion Ling 2.6
Inclusion · inclusionai/ling-2.6-1t
🔒 19 runs at #20 ▲ best #2
★☆☆☆☆ 9.1 0 0 20 7 Budget No endpoint

Free options (comparison)

# Model Stars Overall Speed Reliability CQ ? CQ — Conversation Quality
Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model.
Value TTFT p50 Tier Badges
1 ▲1
FREE LLM
OpenRouter · free-llm-cycler
★★★★☆ 76.8 50 95 86 72 2.06s Free Free option
2 ▼1
Free LLM Router
Open Router · openrouter/free
▲ best #1
★★★☆☆ 63.7 35 91 70 58 2.58s Free Free option Not recommended for live co-hosts

/ = rank change since the previous run · = no change · NEW = first appearance

Probes are scripted OpenRouter checks (not live mixer sessions). Pass/fail + timings only — no reply transcripts stored. Scores reflect a sampled run (a few attempts per model at a non-zero temperature), so small differences between runs are normal — treat close scores as ties, not regressions.