PR #534 agent changes, tested one at a time
As written, PR #534 makes answers worse. Against production it wins only 36% of blind head-to-heads (38 wins, 69 losses; p = 0.004), mainly because answers become less complete. It is 16% cheaper per turn, but the first text arrives about 5 s later.
No single change causes this. Tested alone, every change except the evidence byte caps is quality-neutral. The regression comes from their combined effect on how much evidence reaches the model: 8 passages per search instead of 20, plus the 24/48 KiB caps, cuts delivered evidence from about 165 KiB to about 60 KiB per turn.
The fix is small and measured. Keep every PR change except two: deliver 20 passages instead of 8, and remove the byte caps. That version is at parity with production (54%, 51–44), costs 6% less, adds only 1.4 s to first text, and keeps the PR's real wins: working expansion, catalog-filtered search, dedup and the budget.
Recommendations
- Do not merge as-is. Change
RETRIEVAL_DELIVERED8 → 20, with hybrid at 25 per branch, and remove the 24/48 KiB evidence caps (or raise them above about 200 KiB so they only catch outliers). This is the only configuration that reached parity (54%), and it is still 6% cheaper than production. - Keep hybrid search, lookup_catalog with author, book and genre filters, grounded expansion, dedup and memoization, the no-minimum-search prompt line, prompt v28, and keeping tool history on follow-ups. Each is quality-neutral or better and fixes a real gap. Catalog filters win 64% on named-author questions, and expansion is broken in production today.
- Fix the budget cut-off before release. When the 8-call budget runs out,
prepareStepremoves all tools, but Gemini keeps emittingsearchcalls, which surface asAI_NoSuchToolError. Parallel calls within one step can also exceed the cap and silently get empty results. Return an explicit "budget exhausted" tool result instead of removing the tools. - Hold the date filter (
deathYearAH) until the backfill has run and been verified. On the live index today, any date-filtered search fails with HTTP 400 ("attribute not found"). The model also used the filter in only 1 of 92 real questions, including 0 of 12 "salaf / early scholars" questions, so it is low value until the prompt teaches when to use it. - Expect a latency cost. The PR adds sequential tool rounds (2.4 vs 1.2 before the answer). Consider streaming a status earlier, or letting the first round issue more parallel searches.
- Re-run this harness on the revised PR (
evals/pr534-ablation, one command per arm) before merge, and extend it to the Quran, hadith and entity routes, which also change prompts but were not tested here. - Ops: production Langfuse tracing appears to have stopped on 16 Sep, 19:39 UTC; no chat traces exist after that. Please check the exporter.
Each atomic change, measured alone
Each row turns on one PR change on top of production (main code + prompt v24) and replays the same real questions. Costs and times are on the 80 first-turn questions unless noted. Quality is the order-collapsed blind win rate against production over decided pairs (95% CI); the production-vs-production null control landed at 48%. Green or red means the CI excludes 50%; grey means no measurable difference. Rows at the bottom are combinations.
| Change | Cost / turn vs prod | First text (p50) vs prod | Total time (p50) vs prod | Answer quality vs prod (blind win rate, decided pairs, all judged slices) | Other effects | Recommendation |
|---|---|---|---|---|---|---|
| Always hybrid search Vector + BM25 (25 each) fused with RRF, then rerank → 20; the model no longer picks semantic/keyword. |
+10% $0.0585 |
+0.7 s 14.2 s |
+0.2 s | 45% 15W / 18L / 7T · CI 30%–62% · no clear difference | Search latency +265 ms (1.68 s vs 1.41 s p50). Evidence +16% because BM25 favours long passages. Quality 45%, n.s. (Arabic 38%, English 53%). | Take Quality-neutral; needed to remove the unreliable mode choice. Its cost is absorbed in the recommended version. |
| 8 passages per search (was 20) Reranker returns 8 instead of 20. |
-24% $0.0404 |
+0.8 s 14.4 s |
+0.6 s | 53% 17W / 15L / 8T · CI 36%–69% · no clear difference | Largest cost cut alone (−24%). Completeness already trends worse alone (9–16). With the other changes it drives the regression: the PR at 8 passages scores 36–38%, and at 20 it scores 54%. | Drop Keep 20. Evidence volume is what moves answer quality (same finding as the Aug rerank study). |
| lookup_catalog + source filters Resolve author, book or genre names to IDs; search accepts authorIds, bookIds and categoryIds. |
+11% $0.0472 |
+3.0 s 15.6 s |
-1.0 s | 64% 9W / 5L / 6T · CI 39%–84% · no clear difference | On random questions: +2% cost, 49% quality. Filters are applied on only 2 of 20 named questions without prompt v28, and 9 of 20 with it. The alias “ابن قيم الجوزية” missed; the model recovered. | Take Wins 64% on named-author and named-book questions; ship together with prompt v28. |
| Grounded expansion The model sees version_id and sequence_number; expand is validated against retrieved passages; ±2 neighbours (was ±5). |
+9% $0.0579 |
+1.2 s 14.8 s |
+1.2 s | 57% 38W / 29L / 29T · CI 45%–68% · no clear difference | Production expand is broken: the model never sees coordinates, guesses them and gets 0 rows (0 of 12 calls across production-code arms succeeded). With the PR, 11/14 succeed and 3 guessed coordinates are rejected. +9% cost. | Take Fixes a dead tool; quality 57% (38–29), the best single change. |
| Dedup + memoization Identical calls are cached; already-delivered passages are dropped and excluded (NotIn) from later searches. |
-6% $0.0500 |
+1.3 s 14.9 s |
+2.6 s | 51% 25W / 24L / 7T · CI 37%–64% · no clear difference | −6% cost. The 3.4% unretrieved citations here were malformed IDs (truncated UUIDs, :page suffixes), concentrated in one run, not a dedup effect. | Take Quality-neutral (51%), small saving, prevents repeated evidence. |
| Retrieval call budget At most 8 retrieval calls; stop after 2 empty calls; tools off at step 18. |
+5% $0.0561 |
+1.1 s 14.6 s |
-0.1 s | 53% 18W / 16L / 6T · CI 37%–69% · no clear difference | Rarely binds on production code (4/80 turns). Inside the PR it binds on 7/80 (13/80 with byte caps), and then Gemini still calls search → AI_NoSuchToolError. | Fix, then take Quality-neutral (53%). Fix the tool-removal behaviour first (see findings). |
| Evidence byte caps 24 KiB per tool result, 48 KiB per model step; the newest evidence is kept and older evidence is dropped from context. |
-8% $0.0491 |
+2.8 s 16.3 s |
+0.1 s | 29% 13W / 32L / 11T · CI 18%–43% · worse | The model re-searches for evidence that was dropped: searches 5.3 vs 3.4, budget hits 13/80, first text +2.7 s. Judges cite thinner, mis-cited and invented quotations. On Arabic it is even more expensive (+7%). | Drop The only atomic change that is clearly worse: 29% (13–32), p = 0.007. |
| Keep tool history on follow-ups pruneMessages toolCalls: "none"; earlier turns' evidence stays in context (needed for follow-up expansion). |
+12% $0.0364 |
-3.9 s 5.6 s |
-0.5 s | 64% 7W / 4L / 5T · CI 35%–85% · no clear difference | Follow-up first text is 3.8 s faster (the model reuses evidence instead of searching). +12% cost from larger inputs. Only 16 conversations. | Take 64% on follow-ups (7–4, small n); faster; enables grounded expansion. |
| Prompt: no minimum number of searches Removes “minimum of two initial searches for simple questions, four for complex” from v24. |
-16% $0.0449 |
+0.2 s 13.8 s |
-0.3 s | 50% 13W / 13L / 14T · CI 32%–68% · no clear difference | Searches 2.5 vs 3.4; −16% cost with no change in first-text time. | Take Free saving: quality 50% (13–13–14). |
| Combined: PR as written Every change above, 8 passages, byte caps, prompt v28. |
-16% $0.0448 |
+5.0 s 18.6 s |
+1.8 s | 36% 38W / 69L / 21T · CI 27%–45% · worse | Tool rounds before the answer 2.4 vs 1.2. Cited IDs 8.9 vs 16.7. Completeness 27–77. 9 turns with tool errors. | Don’t merge as-is 36% (38–69), p = 0.004; Arabic 31%. |
| Combined: PR code, old prompt v24 Isolates the prompt: the same code with production prompt v24. |
-16% $0.0447 |
+4.5 s 18.0 s |
+1.7 s | 34% 32W / 63L / 33T · CI 25%–44% · worse | v28 is better than v24 on the same code: 61% (63–40), p = 0.03. The prompt is not the problem. | — 34% (32–63): the regression lives in the code changes. |
| Combined: PR without byte caps Leave-one-out: everything except the byte caps. |
-28% $0.0382 |
+2.9 s 16.5 s |
+2.5 s | 38% 43W / 70L / 15T · CI 30%–47% · worse | Cheapest arm (−28%). 0 tool errors. Still thin evidence at 8 passages. | Not enough 38% (43–70), p = 0.014. Removing the caps alone does not fix quality. |
| Combined: PR, no byte caps, 20 passages Recommended: every PR change + v28, without byte caps, delivering 20 passages (hybrid 25/branch). |
-6% $0.0501 |
+1.5 s 15.0 s |
+1.4 s | 54% 51W / 44L / 33T · CI 44%–63% · no clear difference | Evidence 128 KiB/turn. Named-source questions 62%, Arabic 59%, chronology 40% (n = 12). Completeness tied 48–54. | Ship this Parity at 54% (51–44, CI 44–63), −6% cost, +1.4 s to first text. |
Bugs and risks found while testing
A direct Turbopuffer query with the PR's author_death_year filter returns HTTP 400: filter error in key author_death_year: attribute not found. Until the manual backfill runs and is verified, every date-filtered search call is a tool error. The PR body says release order is backfill first; this confirms it is a hard gate, not a nicety. The model also rarely uses the filter: once in 92 real questions, and never on 12 real “salaf / early scholars / المتقدمين” questions.
When the retrieval budget is exhausted, prepareRetrievalStep returns activeTools: [] with toolChoice: "none", but Gemini still emits search calls, which surface as AI_NoSuchToolError (2–3 per affected turn in replays). The budget is also checked per call, not per step: a step with 10 parallel searches runs 8 and silently returns [] for the rest. Suggest keeping the tools active and returning an explicit “retrieval budget exhausted — answer from the evidence above” result.
With older evidence dropped from context, the model searches again for what it lost: 5.3 vs 3.4 searches per turn, and the budget is hit in 13 of 80 turns vs 1 of 80. The saving from shorter inputs is eaten by the extra rounds (on Arabic questions cost goes *up* 7%), and first text is 2.7 s slower.
readEvidence in retrieval-evidence.ts parses tool outputs with the full chunk schema and returns [] on any mismatch. For example, results without version_id are erased from the model's context with no log. The PR's own outputs pass, but any future format change or odd historical message would vanish silently. The harness hit this directly. Suggest logging and passing through unparseable results instead of dropping them.
On main, formatChunkForModel hides version_id and there is no sequence_number, so the model guesses coordinates and gets 0 rows: none of the 12 expand calls made by production-code arms returned anything. The PR's grounded expansion fixes this (11 of 14 succeed; 3 invented coordinates were correctly rejected).
handlers/chat/index.ts changes pruneMessages({ toolCalls }) from before-last-2-messages to none. This is required for follow-up expansion. It makes follow-up first text 3.8 s faster but costs 12% more per follow-up. Bigger histories also move the Gemini implicit-cache prefix; check prod cache-hit rates after release.
The newest production chat-message trace is from 2026-09-16 19:39 UTC, and nothing arrived in the 10 days since. This harness sampled 6–16 Sep for that reason. The Langfuse exporter or keys likely broke around a deploy that day.
Full PR vs production
The PR as written loses to production on completeness, not tone or language. Judges repeatedly note fewer salaf sayings, scholars or verses than the question asked for, whole sections without citations, and claims cited to passages that do not contain them. That is what an evidence-starved model does: it sees about 60 KiB of evidence per turn instead of about 165 KiB, and cites 8.9 passages instead of 16.7.
Speed: total turn time is about the same (answer generation dominates), but the model makes 2.4 sequential tool rounds before answering instead of 1.2, so first text moves from 13.6 s to 18.6 s (median). The recommended version makes 1.4 rounds and lands at 15.0 s.
Cost: −16% as written. The recommended version is −6%: most of the saving came from starving the model, which is what cost the quality.
First-turn questions (80: 40 Arabic, 40 Latin-script)
| Arm | n | $/turn | First text s (p50) | Total s (p50) | Steps | Searches | Expands | Lookups | Input tok | Evidence | Words | Cited ids | Unretrieved cites | Turns w/ tool error |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Production (main) | 80 | $0.0533 | 13.6 | 29.7 | 2.2 | 3.4 | 0.03 | 0.00 | 48k | 161 KiB | 500 | 16.7 | 0.9% | 0 |
| Production replicate | 80 | $0.0500 | 13.8 | 30.6 | 2.1 | 3.2 | 0.01 | 0.00 | 44k | 143 KiB | 509 | 16.9 | 0.5% | 0 |
| PR #534 as written | 80 | $0.0448 | 18.6 | 31.6 | 3.4 | 3.6 | 0.19 | 0.19 | 35k | 58 KiB | 445 | 8.9 | 2.4% | 9 |
| PR code, prompt v24 | 80 | $0.0447 | 18.0 | 31.5 | 3.0 | 4.6 | 0.06 | 0.14 | 30k | 68 KiB | 440 | 10.1 | 0.4% | 12 |
| PR without byte caps | 80 | $0.0382 | 16.5 | 32.2 | 2.9 | 2.9 | 0.07 | 0.17 | 36k | 56 KiB | 447 | 11.5 | 0.5% | 0 |
| PR, no byte caps, 20 passages | 80 | $0.0501 | 15.0 | 31.1 | 2.4 | 2.5 | 0.05 | 0.16 | 47k | 125 KiB | 492 | 16.3 | 0.5% | 0 |
| Hybrid search only | 80 | $0.0585 | 14.2 | 29.9 | 2.1 | 3.3 | 0.00 | 0.00 | 52k | 188 KiB | 498 | 17.8 | 1.4% | 0 |
| 8 passages only | 80 | $0.0404 | 14.4 | 30.4 | 2.5 | 3.7 | 0.03 | 0.00 | 34k | 74 KiB | 444 | 12.9 | 0.5% | 0 |
| Catalog lookup + source filters | 80 | $0.0544 | 14.7 | 29.8 | 2.3 | 3.3 | 0.01 | 0.21 | 50k | 164 KiB | 476 | 16.7 | 1.0% | 0 |
| Grounded expansion | 80 | $0.0579 | 14.8 | 30.9 | 2.2 | 3.6 | 0.03 | 0.00 | 55k | 174 KiB | 540 | 16.5 | 0.8% | 0 |
| Dedup + memo | 79 | $0.0500 | 14.9 | 32.3 | 2.2 | 3.4 | 0.01 | 0.00 | 44k | 129 KiB | 511 | 17.3 | 3.4% | 0 |
| Retrieval call budget | 80 | $0.0561 | 14.6 | 29.6 | 2.2 | 3.5 | 0.03 | 0.00 | 53k | 165 KiB | 514 | 16.7 | 0.7% | 0 |
| Evidence byte caps | 80 | $0.0491 | 16.3 | 29.8 | 3.2 | 5.3 | 0.11 | 0.00 | 33k | 93 KiB | 440 | 10.3 | 2.9% | 0 |
| No minimum searches (prompt) | 80 | $0.0449 | 13.8 | 29.4 | 2.4 | 2.5 | 0.04 | 0.00 | 44k | 96 KiB | 498 | 15.6 | 0.9% | 1 |
Quality by language
Arabic
| Contrast | Win rate of left arm (decided pairs) | |
|---|---|---|
| full vs base | 31% 10W / 22L / 8T · CI 18%–49% · worse | |
| pr_v24 vs base | 26% 8W / 23L / 9T · CI 14%–43% · worse | |
| full_nobytes vs base | 31% 11W / 24L / 5T · CI 19%–48% · worse | |
| full_20 vs base | 59% 16W / 11L / 13T · CI 41%–75% · no clear difference | |
| full vs pr_v24 | 66% 23W / 12L / 5T · CI 49%–79% · no clear difference | |
| size8 vs base | 44% 8W / 10L / 2T · CI 25%–66% · no clear difference | |
| min_search vs base | 50% 6W / 6L / 8T · CI 25%–75% · no clear difference | |
| hybrid vs base | 38% 6W / 10L / 4T · CI 18%–61% · no clear difference | |
| bytes vs base | 29% 5W / 12L / 3T · CI 13%–53% · no clear difference | |
| base2 vs base | 47% 14W / 16L / 10T · CI 30%–64% · no clear difference |
Latin-script
| Contrast | Win rate of left arm (decided pairs) | |
|---|---|---|
| full vs base | 39% 14W / 22L / 4T · CI 25%–55% · no clear difference | |
| pr_v24 vs base | 43% 13W / 17L / 10T · CI 27%–61% · no clear difference | |
| full_nobytes vs base | 36% 12W / 21L / 7T · CI 22%–53% · no clear difference | |
| full_20 vs base | 52% 16W / 15L / 9T · CI 35%–68% · no clear difference | |
| full vs pr_v24 | 52% 16W / 15L / 9T · CI 35%–68% · no clear difference | |
| size8 vs base | 64% 9W / 5L / 6T · CI 39%–84% · no clear difference | |
| min_search vs base | 50% 7W / 7L / 6T · CI 27%–73% · no clear difference | |
| hybrid vs base | 53% 9W / 8L / 3T · CI 31%–74% · no clear difference | |
| bytes vs base | 27% 4W / 11L / 5T · CI 11%–52% · no clear difference | |
| base2 vs base | 50% 16W / 16L / 8T · CI 34%–66% · no clear difference |
Targeted slices
Named authors and books: this is the slice the catalog tool is for. With prompt v28 the model resolves the name and filters in 9 of 20 questions, and the catalog change alone wins 64% there. The PR as written still loses (35%) because of thin evidence; the recommended version wins 62% (8–5), directionally better than production.
Chronology: real “salaf / early vs later scholars” questions never triggered the date filter, so they measure evidence volume more than chronology. The PR as written is slowest here (27.7 s to first text). The recommended version is at 40% on only 12 questions (CI 17–69), so no conclusion either way.
Follow-ups: production English follow-ups are rare (1 in about 250 multi-turn traces), so this set of 16 is 8 Arabic and 8 Latin-script, including French and Indonesian. Keeping tool history is the useful change here; everything is within noise at this size.
Named author or book (20)
| Arm | n | $/turn | First text s (p50) | Total s (p50) | Steps | Searches | Expands | Lookups | Input tok | Evidence | Words | Cited ids | Unretrieved cites | Turns w/ tool error |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Production (main) | 20 | $0.0423 | 12.6 | 28.2 | 2.1 | 2.8 | 0.00 | 0.00 | 36k | 110 KiB | 416 | 11.6 | 0.4% | 0 |
| PR #534 as written | 20 | $0.0412 | 19.1 | 32.7 | 3.6 | 2.9 | 0.00 | 0.60 | 34k | 47 KiB | 357 | 7.7 | 0.7% | 0 |
| PR code, prompt v24 | 20 | $0.0378 | 15.0 | 29.3 | 2.8 | 3.7 | 0.00 | 0.35 | 24k | 60 KiB | 405 | 8.6 | 0.6% | 0 |
| PR without byte caps | 20 | $0.0302 | 15.3 | 27.5 | 2.8 | 2.1 | 0.00 | 0.60 | 24k | 40 KiB | 365 | 9.2 | 0.0% | 0 |
| PR, no byte caps, 20 passages | 20 | $0.0432 | 15.9 | 31.5 | 2.6 | 1.9 | 0.00 | 0.50 | 41k | 97 KiB | 418 | 14.1 | 0.4% | 0 |
| Catalog lookup + source filters | 20 | $0.0472 | 15.6 | 27.3 | 2.3 | 3.0 | 0.00 | 0.30 | 42k | 162 KiB | 422 | 13.8 | 0.4% | 0 |
| Contrast | Win rate of left arm (decided pairs) | |
|---|---|---|
| full vs base | 35% 6W / 11L / 3T · CI 17%–59% · no clear difference | |
| full_nobytes vs base | 47% 8W / 9L / 3T · CI 26%–69% · no clear difference | |
| full_20 vs base | 62% 8W / 5L / 7T · CI 36%–82% · no clear difference | |
| pr_v24 vs base | 27% 3W / 8L / 9T · CI 10%–57% · no clear difference | |
| catalog vs base | 64% 9W / 5L / 6T · CI 39%–84% · no clear difference | |
| full vs pr_v24 | 64% 9W / 5L / 6T · CI 39%–84% · no clear difference |
Chronology / “salaf”, “early scholars” (12)
| Arm | n | $/turn | First text s (p50) | Total s (p50) | Steps | Searches | Expands | Lookups | Input tok | Evidence | Words | Cited ids | Unretrieved cites | Turns w/ tool error |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Production (main) | 12 | $0.0587 | 13.8 | 38.6 | 2.2 | 4.2 | 0.00 | 0.00 | 53k | 182 KiB | 576 | 19.7 | 0.4% | 0 |
| PR #534 as written | 12 | $0.0548 | 27.7 | 43.2 | 3.8 | 4.7 | 0.00 | 0.00 | 42k | 105 KiB | 467 | 10.2 | 0.8% | 1 |
| PR code, prompt v24 | 12 | $0.0486 | 18.8 | 40.5 | 3.1 | 6.2 | 0.00 | 0.00 | 32k | 90 KiB | 460 | 10.6 | 0.8% | 4 |
| PR without byte caps | 12 | $0.0472 | 21.7 | 43.2 | 3.1 | 3.8 | 0.00 | 0.00 | 46k | 80 KiB | 554 | 14.7 | 0.6% | 1 |
| PR, no byte caps, 20 passages | 12 | $0.0502 | 15.6 | 30.7 | 2.2 | 2.7 | 0.00 | 0.00 | 49k | 113 KiB | 550 | 16.3 | 0.5% | 0 |
| Catalog lookup + source filters | 12 | $0.0572 | 17.0 | 34.5 | 2.2 | 3.4 | 0.00 | 0.00 | 51k | 188 KiB | 576 | 17.6 | 0.9% | 0 |
| Contrast | Win rate of left arm (decided pairs) | |
|---|---|---|
| full vs base | 27% 3W / 8L / 1T · CI 10%–57% · no clear difference | |
| full_nobytes vs base | 42% 5W / 7L / 0T · CI 19%–68% · no clear difference | |
| full_20 vs base | 40% 4W / 6L / 2T · CI 17%–69% · no clear difference | |
| pr_v24 vs base | 20% 2W / 8L / 2T · CI 6%–51% · no clear difference | |
| catalog vs base | 50% 4W / 4L / 4T · CI 22%–78% · no clear difference | |
| full vs pr_v24 | 70% 7W / 3L / 2T · CI 40%–89% · no clear difference |
Follow-up turns (16 two-turn conversations; second turn measured)
| Arm | n | $/turn | First text s (p50) | Total s (p50) | Steps | Searches | Expands | Lookups | Input tok | Evidence | Words | Cited ids | Unretrieved cites | Turns w/ tool error |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Production (main) | 16 | $0.0324 | 9.4 | 19.6 | 2.0 | 1.6 | 0.00 | 0.00 | 32k | 49 KiB | 420 | 9.9 | 3.8% | 0 |
| PR #534 as written | 16 | $0.0330 | 15.7 | 26.4 | 2.6 | 2.4 | 0.00 | 0.44 | 28k | 39 KiB | 367 | 6.7 | 0.0% | 0 |
| PR code, prompt v24 | 16 | $0.0455 | 17.5 | 29.6 | 3.2 | 3.6 | 0.00 | 0.44 | 37k | 43 KiB | 473 | 7.8 | 0.8% | 2 |
| PR without byte caps | 16 | $0.0350 | 13.8 | 26.6 | 2.8 | 2.4 | 0.00 | 0.44 | 37k | 42 KiB | 393 | 9.5 | 0.0% | 0 |
| PR, no byte caps, 20 passages | 16 | $0.0397 | 12.4 | 23.8 | 2.4 | 1.8 | 0.00 | 0.25 | 45k | 56 KiB | 363 | 10.1 | 0.0% | 0 |
| Grounded expansion | 16 | $0.0433 | 12.3 | 23.3 | 2.1 | 2.5 | 0.00 | 0.00 | 45k | 77 KiB | 456 | 10.2 | 2.5% | 0 |
| Keep tool history | 16 | $0.0364 | 5.6 | 19.0 | 2.1 | 1.9 | 0.00 | 0.00 | 42k | 51 KiB | 391 | 9.2 | 4.8% | 0 |
| Dedup + memo | 16 | $0.0415 | 14.2 | 27.2 | 2.0 | 2.6 | 0.00 | 0.00 | 41k | 90 KiB | 460 | 12.0 | 5.2% | 0 |
| Evidence byte caps | 16 | $0.0317 | 10.9 | 22.6 | 2.4 | 2.6 | 0.00 | 0.00 | 28k | 47 KiB | 414 | 5.8 | 2.2% | 0 |
| Contrast | Win rate of left arm (decided pairs) | |
|---|---|---|
| full vs base | 45% 5W / 6L / 5T · CI 21%–72% · no clear difference | |
| full_nobytes vs base | 44% 7W / 9L / 0T · CI 23%–67% · no clear difference | |
| full_20 vs base | 50% 7W / 7L / 2T · CI 27%–73% · no clear difference | |
| pr_v24 vs base | 46% 6W / 7L / 3T · CI 23%–71% · no clear difference | |
| expand vs base | 50% 4W / 4L / 8T · CI 22%–78% · no clear difference | |
| history vs base | 64% 7W / 4L / 5T · CI 35%–85% · no clear difference | |
| session vs base | 44% 7W / 9L / 0T · CI 23%–67% · no clear difference | |
| bytes vs base | 31% 4W / 9L / 3T · CI 13%–58% · no clear difference | |
| full vs pr_v24 | 62% 8W / 5L / 3T · CI 36%–82% · no clear difference |
Why judges preferred one side
- Wins when production cites fabricated or unsupplied passage IDs, leaks raw citation markup, or ignores an explicit constraint such as “contemporary scholars” or “salaf” (about 16 pairs).
- Loses on coverage: fewer sayings, scholars or stories than asked (about 11 pairs), whole uncited sections (about 8), claims cited to passages that do not contain them (about 9), and softened or misstated positions (about 7).
- Loses on claims attached to passages that do not contain them: wrong author (al-Majd vs Shaykh al-Islam, Sayyid vs Muhammad Qutb) or bibliography lists used as doctrinal support (20+ pairs).
- Thinner answers (about 12), fabricated passages cited for central claims (about 11), and invented or spliced quotation text (about 7). These are the symptoms of evidence silently dropping out of context between steps.
- Judges call it about even; repeated fixtures often flip winner.
- It wins when production cites fabricated passages or misses a requested element.
- It loses on individual mis-citations and occasional dropped views.
- Completeness is tied (48–54), which was the gap that sank the PR as written.
Method and caveats
- Data: real production chat turns pulled from Langfuse, 6–16 Sep 2026 (prompt v24, gemini-3.7-flash), random sample, excluding greetings and custom instructions. 80 first-turn questions, 40 Arabic and 40 Latin-script (mostly English, some French, Indonesian and similar). 20 named author/book questions and 12 chronology questions, found by pattern. 16 two-turn conversations, where both turns are regenerated per arm so history pruning and expansion behave as in production.
- Replay: the real pipeline pieces — Langfuse prompts v24/v28 with personalization, the production Turbopuffer namespace, Cohere rerank v4 pro, the citation middleware, 20-step cap and 32k output cap — on Vertex gemini-3.7-flash. Production routes the same model through the Cloudflare AI Gateway; no gateway keys were available locally. The tools are re-implemented behind one flag per PR change, mirroring main and the PR's
tools.ts,retrieval-session.ts,retrieval-evidence.ts,retrieval-step.tsandretrieval-filters.tsand importing the PR's own helpers. The date filter is simulated through the equivalent author-ID set, because the live index lacks the attribute. - Quality: blind pairwise judging by Claude Opus 5.5, 5 pairs per agent, both presentation orders in separate batches, with the passages behind every citation attached and citation IDs blinded. A pair counts as decided only when both orders agree. Win rate = wins ÷ decided, with a 95% Wilson CI and a two-sided sign test. The workflow judged 20+20 questions per contrast first and extended to all 80 only where round 1 was undecided. 2,610 verdicts; null control (production vs production replicate) 48%.
- Cost: token usage × $0.75 / $3.75 per M input/output (cached input 10%) + $0.0025 per Cohere rerank. Embeddings and Turbopuffer are excluded (negligible). Replays see little implicit caching, so absolute costs are slightly above production's; the relative differences are what matter.
- Caveats: one run per arm per question, so n = 80 gives about ±10 points on quality. Single-change arms decided on 30–60 pairs can only detect large effects. The Quran, hadith and entity routes (which also get new prompts and tools) and the UI/playground changes are not covered. Two harness confounds in the byte-caps-only arm were found and fixed before judging: projection needs
version_id, and it stripsmodefrom history. Latency was measured from a laptop to Vertexglobal; compare arms to each other, not to production dashboards. - Reproduce:
evals/pr534-ablation/on the PR branch:collect.ts→run.ts --arms …→prepare-batches.ts→judge.workflow.js→extract.ts→analyze.ts→build-report.ts.
Prepared for the Qaf team. Internal; not indexed.