Skip to main content
Blog

Peer review, engineering, and adversarial AI.

Insights and tutorials from the tellodb team. We write about how we build the pipeline, how teams use it, and what we're learning about making models disagree productively.

US labs vs Alibaba vs Meta in the same debate

Why seat Claude or GPT, Qwen3.8 Max, and Muse Spark 1.2 together: correlated failure modes, lab map, Delphi silent first round, and links to false consensus research on RigorDesk.

Read →·by Sharjeel Abbas

What the Research Actually Says About Multi-Agent LLM Debate

Multi-agent debate mostly fails to beat simple self-consistency at matched compute. Here is what the 2023-2026 literature actually found, why homogeneous debate collapses into sycophancy, and the one variable that survives every negative result.

Read →·by Sharjeel Abbas

See What the First AI Missed

Learn how to get a useful AI second opinion that finds missing assumptions, weak evidence, counterexamples, risks, and decision-changing errors.

Read →·by Sharjeel Abbas

How to Evaluate AI Agents in 2026

A production guide to evaluating AI agents with deterministic tests, calibrated LLM judges, agent judges, human review, trajectory analysis, and risk-based release gates.

Read →·by Sharjeel Abbas

Nine AI Judges, Two Real Opinions

Why nine LLM judges may provide only two effectively independent opinions—and how to design multi-model review that resists correlated errors and false consensus.

Read →·by Sharjeel Abbas

LLM Debate vs. Single-Prompt Bias

A deep dive into why single models fail at logic, and how adversarial review fixes it. LLM debate, multi-model critique, and structured disagreement.

Read →·by Sharjeel Abbas