Podnami

Best performance on Podnami

LLM Rankings

Models from Settings → AI Selection, scored for live co-hosts: speed (TTFT), reliability, conversation quality (CQ), and credit value. High speed alone does not mean high CQ.

Last completed run: 2026-08-14T01:17:09.334537+00:00 · 19 models

Click a column header to sort.

Ranked models

# Model Stars Overall Speed Reliability CQ ? CQ — Conversation Quality
Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model.
Value TTFT p50 Tier Badges
1 ▲1
Gemini 2.5 Lite
Google · google/gemini-2.5-flash-lite
★★★★★ 99.4 100 100 99 100 0.45s Budget Best value
2 ▲11
Inclusion Ling 2.6 Flash
Ling · inclusionai/ling-2.6-flash
★★★★★ 97.8 95 100 99 97 0.54s Budget
3 ▲3
Google Gemini Flash Lite 3.5
Google · google/gemini-3.5-flash-lite
★★★★★ 95.7 100 100 99 75 0.42s Standard
4 ▲11
Mistral Nemo
Mistral · mistralai/mistral-nemo
★★★★★ 95.7 95 100 94 96 0.45s Budget Not recommended for live co-hosts
5 ▲9
GPT-3.5 Turbo
OpenAI · openai/gpt-3.5-turbo
▲ best #4
★★★★★ 94.0 100 100 95 74 0.31s Standard
6 ▲2
Amazon Nova Lite V1
Amazon · amazon/nova-lite-v1
▲ best #1
★★★★★ 93.3 95 100 89 94 0.47s Budget
7 ▼6
DeepSeek Chat
DeepSeek · deepseek/deepseek-chat
▲ best #1
★★★★★ 93.1 80 100 99 91 1.57s Budget
8 ▲1
DeepseekV32
DeepSeek · deepseek/deepseek-v3.2
▲ best #3
★★★★★ 91.4 80 100 95 89 1.23s Budget
9 ▲1
Ministral 14b-2512
Ministral 14b-2512 · mistralai/ministral-14b-2512
▲ best #2
★★★★★ 91.4 80 100 95 89 0.44s Budget
10 ▼3
Meta Llamma Scout
Meta Llamma Scout · meta-llama/llama-4-scout
▲ best #2
★★★★★ 91.4 80 100 95 89 0.31s Budget
11
GPT-4 Turbo
OpenAI · openai/gpt-4-turbo
🔒 2 runs at #11
★★★★★ 91.3 95 100 99 54 1.42s Premium
12 ▼9
GPT-4o Mini
OpenAI · openai/gpt-4o-mini
▲ best #3
★★★★★ 89.7 80 100 99 68 0.51s Standard
13 ▼9
Claude 3 Haiku
Anthropic · anthropic/claude-3-haiku
▲ best #4
★★★★★ 89.7 80 100 99 68 0.34s Standard
14 ▲2
Gemma 4 31B IT
Google · google/gemma-4-31b-it
▲ best #3
★★★★★ 89.7 80 100 91 88 0.34s Budget
15 ▼3
DeepSeek V4 Flash
DeepSeek · deepseek/deepseek-v4-flash
▲ best #12
★★★★★ 86.1 65 95 96 82 1.63s Budget
16 ▼11
Claude Haiku 4.5
Anthropic · anthropic/claude-haiku-4.5
▲ best #5
★★★★★ 85.2 65 100 99 63 0.89s Standard
17
Step 3.5 Flash
StepFun · stepfun/step-3.5-flash
🔒 8 runs at #17
★★★☆☆ 58.0 45 62 65 55 2.66s Budget

Free options (comparison)

# Model Stars Overall Speed Reliability CQ ? CQ — Conversation Quality
Suitability for Podnami live co-hosts: tool calling, no speaker-label leaks, handoffs, short memory, spoken brevity, an in-character chat check, plus whether the model is a non-thinking chat-oriented model vs a utility/doc or heavy-reasoning model.
Value TTFT p50 Tier Badges
1
FREE LLM
OpenRouter · free-llm-cycler
🔒 3 runs at #1
★★★★☆ 70.3 45 100 73 66 1.78s Free Free option Not recommended for live co-hosts
2
Free LLM Router
Open Router · openrouter/free
🔒 3 runs at #2 ▲ best #1
★★★☆☆ 66.8 45 90 70 63 2.46s Free Free option Not recommended for live co-hosts

/ = rank change since the previous run · = no change · NEW = first appearance

Probes are scripted OpenRouter checks (not live mixer sessions). Pass/fail + timings only — no reply transcripts stored. Scores reflect a sampled run (a few attempts per model at a non-zero temperature), so small differences between runs are normal — treat close scores as ties, not regressions.