disclosure-bureau

discadmin/disclosure-bureau

Fork 0

Commit graph

Author	SHA1	Message	Date
Luiz Gustavo	7826710051	W4: bilingual EN + PT-BR Investigation Bureau (CLAUDE.md §3 contract) Some checks failed CI / Web — typecheck + lint + build (push) Failing after 41s Details CI / Scripts — Python smoke (push) Failing after 4s Details CI / Web — npm audit (push) Failing after 26s Details CI / Retrieval — golden set (Recall@5 + MRR) (push) Failing after 4s Details User flagged that the bureau was emitting English-only output, violating the project's bilingual rule. Every narrative field now ships in both languages: stored in sibling DB columns + rendered as adjacent markdown sections per CLAUDE.md §3. Migration 0007 (apply as supabase_admin): - public.hypotheses +question_pt_br, +position_pt_br, +argument_for_pt_br, +argument_against_pt_br - public.contradictions +topic_pt_br, +notes_pt_br - public.witnesses +access_to_event_pt_br, +bias_notes_pt_br, +verdict_pt_br - public.gaps +description_pt_br, +suggested_next_move_pt_br - public.evidence: unchanged (verbatim_excerpt stays source-language) - JSONB siblings inside contradictions.chunks + gaps.scope handled at runtime (statement_pt_br, title_pt_br, dominant_model_pt_br, why_surprising_pt_br, what_it_implies_pt_br). Detective prompts (all 7) rewritten with explicit bilingual JSON contract: - Output protocol section names every EN field + its _pt_br sibling - "Bilingual is mandatory" warning in the task instruction - Sentinel skip-states unchanged (NO_HYPOTHESES, NO_CONTRADICTIONS, INSUFFICIENT_TESTIMONY, INSUFFICIENT_HYPOTHESIS, NO_OUTLIERS, NO_NEW_EVIDENCE, INSUFFICIENT_ARTEFACTS) - Schneier: parallel arrays — hidden_assumptions[i] matches hidden_assumptions_pt_br[i], lengths must match - Case-Writer: interleaved §1 (EN) / §1 (PT-BR) per act in the body Writer-side validation (all 7 tools): - Reject INSERT if PT-BR sibling missing when EN field is set - Persist both languages atomically in one INSERT (no half-updates) - Markdown renderers write adjacent EN+PT-BR sections in case files (## Argument for (EN) followed by ## Argumento a favor (PT-BR), etc.) Detective parse layer (all 7 detectives): - Coerce both keys from JSON output - "incomplete_bilingual_*" skip reason when either side missing - Defensive: PT-BR fields trimmed + length-capped same as EN Orchestrator propagates question_pt_br + topic_pt_br through job payload to runHolmes / runCaseWriter, mirroring the chat-tool entry point. Web (UI): - /api/jobs/[id] hydrates _pt_br siblings from pg - job-status-poller HypothesisCard: PT-BR primary, EN in <details> fallback when both exist - ContradictionCard: PT-BR statement primary + secondary EN quote - WitnessCard: PT-BR verdict primary + secondary EN quote, panels in PT - GapCard: PT-BR title/why/implies primary - /bureau hub: SELECTs both columns, renders PT-BR primary - /h/[id]: ArgumentPanel renders PT-BR primary with collapsible EN fallback when both exist - BureauSnapshot homepage: position_pt_br / topic_pt_br / verdict_pt_br primary - DocBureauPanel /d/[doc]: same primary-PT-BR pattern - New web/lib/i18n/pick.ts helper (unused yet by chat/agents — kept for future locale-driven switching when both languages are equally full; current rule is PT-BR-first since the user is brasileiro) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-05-24 12:02:59 -03:00
Luiz Gustavo	25f19aee63	W3.7 followup: harden Dupin scoping + chunk_id parsing Some checks failed CI / Web — typecheck + lint + build (push) Failing after 32s Details CI / Scripts — Python smoke (push) Failing after 3s Details CI / Web — npm audit (push) Failing after 27s Details CI / Retrieval — golden set (Recall@5 + MRR) (push) Failing after 3s Details Two regressions surfaced in the smoke test that put Dupin from 0/3 contradictions written → 3/3 in the next run. 1. Single-doc scope was too narrow for Dupin's task. Holmes's question about Sandia returned 4 chunks scoped to one doc, but Dupin's terser "topic" form yielded only 1. Solution: Pass-1 tries the requested doc_id; if the head is < 2 chunks, Pass-2 widens to the whole corpus. Audit event carries `scope_widened` so the case-writer can later flag cross-doc contradictions distinctly. The unscoped retry hit 9 chunks and produced 3 contradictions across 3 different docs. 2. Chunk-block header was ambiguous to the model. `--- doc-id/p007#c0042 ---` led Claude to parse `chunk_id` as "p007#c0042" or "p007/c0042" in the JSON output. write_contradiction then refused the FK lookup with "chunk not found". Fix: - Explicit `doc_id:` / `chunk_id:` / `page:` lines per chunk in the rendered block (no slashes/hashes the model can fold). - Defensive normalizeChunkId() in write_contradiction.ts strips any pNNN prefix and keeps only the trailing cNNNN — so the writer is forgiving without losing strictness on the topic + statement validation. Smoke now produces (job 6deddf4b): R-0001 (3 chunks) — Color of the fireball(s) in incident summaries R-0002 (2 chunks) — Geographic confinement of green-fireball sightings R-0003 (3 chunks) — Whether the phenomenon was exclusively green or also red/multicolored R-0003 connects 3 different declassified documents: the Los Alamos conference (exclusively-green category), a retrospective document (red OR green), and Incident 229 (red, blue, yellow — no green). Real cross-doc contradiction, fully cited. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-05-23 21:42:01 -03:00
Luiz Gustavo	5ac53cb3e2	W3.7: Dupin contradiction-scan detective + UI integration Some checks failed CI / Web — typecheck + lint + build (push) Failing after 39s Details CI / Scripts — Python smoke (push) Failing after 4s Details CI / Web — npm audit (push) Failing after 37s Details CI / Retrieval — golden set (Recall@5 + MRR) (push) Failing after 4s Details Adds the third AI detective in the Investigation Bureau runtime: C. Auguste Dupin, who scans a corpus shortlist for pairs (or small groups) of chunks that cannot both be true under any ordinary reading. Runtime: - prompts/dupin.md — discipline (no contradiction without ≥2 distinct chunk_ids; reject same-vocabulary near-misses; FEW high-confidence over MANY weak ones; emit `NO_CONTRADICTIONS` when corpus is silent) - src/detectives/dupin.ts — hybridSearch with k=18 (more chunks than Holmes because contradictions emerge from comparing dispersed claims), strict JSON-array parsing, AT MOST 3 contradictions per call - src/tools/write_contradiction.ts — validates topic + ≥2 positions drawn from ≥2 distinct chunks, resolves chunk_pk via DB lookup (rejects positions citing unknown chunks), INSERTs into public.contradictions + writes case/contradictions/R-NNNN.md - orchestrator: new `contradiction_scan` kind dispatching to runDupin; payload { topic, doc_id?, lang?, context_chunks? } Chat + UI: - request_investigation gains kind=contradiction_scan + topic arg; triggered detective auto-resolves to dupin - chat-bubble inline card renders dupin in orange (#ff8a4d) to distinguish from holmes (cyan) and locard (green) - /jobs/[id] page swaps title + subtitle + tone per detective; "Question" label becomes "Topic" for contradiction_scan - /api/jobs/[id] hydrates public.contradictions when outputs[] surfaces contradiction_ids - job-status-poller renders ContradictionCard: topic + N positions (verbatim statements quoted, stance label optional, link to source chunk) + optional notes panel, with resolution_status badge (open/resolved/irreconcilable) R-NNNN shares the contradiction_id_seq slot with relation per CLAUDE.md naming — same conceptual class (a connection between two pieces of evidence in tension). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-05-23 21:34:04 -03:00

Author

SHA1

Message

Date

Luiz Gustavo

7826710051

W4: bilingual EN + PT-BR Investigation Bureau (CLAUDE.md §3 contract)

CI / Web — typecheck + lint + build (push) Failing after 41s

Details

CI / Scripts — Python smoke (push) Failing after 4s

Details

CI / Web — npm audit (push) Failing after 26s

Details

CI / Retrieval — golden set (Recall@5 + MRR) (push) Failing after 4s

Details

User flagged that the bureau was emitting English-only output, violating
the project's bilingual rule. Every narrative field now ships in both
languages: stored in sibling DB columns + rendered as adjacent markdown
sections per CLAUDE.md §3.

Migration 0007 (apply as supabase_admin):
  - public.hypotheses    +question_pt_br, +position_pt_br,
                         +argument_for_pt_br, +argument_against_pt_br
  - public.contradictions +topic_pt_br, +notes_pt_br
  - public.witnesses     +access_to_event_pt_br, +bias_notes_pt_br,
                         +verdict_pt_br
  - public.gaps          +description_pt_br, +suggested_next_move_pt_br
  - public.evidence: unchanged (verbatim_excerpt stays source-language)
  - JSONB siblings inside contradictions.chunks + gaps.scope handled at
    runtime (statement_pt_br, title_pt_br, dominant_model_pt_br,
    why_surprising_pt_br, what_it_implies_pt_br).

Detective prompts (all 7) rewritten with explicit bilingual JSON contract:
  - Output protocol section names every EN field + its _pt_br sibling
  - "Bilingual is mandatory" warning in the task instruction
  - Sentinel skip-states unchanged (NO_HYPOTHESES, NO_CONTRADICTIONS,
    INSUFFICIENT_TESTIMONY, INSUFFICIENT_HYPOTHESIS, NO_OUTLIERS,
    NO_NEW_EVIDENCE, INSUFFICIENT_ARTEFACTS)
  - Schneier: parallel arrays — hidden_assumptions[i] matches
    hidden_assumptions_pt_br[i], lengths must match
  - Case-Writer: interleaved §1 (EN) / §1 (PT-BR) per act in the body

Writer-side validation (all 7 tools):
  - Reject INSERT if PT-BR sibling missing when EN field is set
  - Persist both languages atomically in one INSERT (no half-updates)
  - Markdown renderers write adjacent EN+PT-BR sections in case files
    (## Argument for (EN) followed by ## Argumento a favor (PT-BR), etc.)

Detective parse layer (all 7 detectives):
  - Coerce both keys from JSON output
  - "incomplete_bilingual_*" skip reason when either side missing
  - Defensive: PT-BR fields trimmed + length-capped same as EN

Orchestrator propagates question_pt_br + topic_pt_br through job payload
to runHolmes / runCaseWriter, mirroring the chat-tool entry point.

Web (UI):
  - /api/jobs/[id] hydrates _pt_br siblings from pg
  - job-status-poller HypothesisCard: PT-BR primary, EN in <details>
    fallback when both exist
  - ContradictionCard: PT-BR statement primary + secondary EN quote
  - WitnessCard: PT-BR verdict primary + secondary EN quote, panels in PT
  - GapCard: PT-BR title/why/implies primary
  - /bureau hub: SELECTs both columns, renders PT-BR primary
  - /h/[id]: ArgumentPanel renders PT-BR primary with collapsible EN
    fallback when both exist
  - BureauSnapshot homepage: position_pt_br / topic_pt_br / verdict_pt_br
    primary
  - DocBureauPanel /d/[doc]: same primary-PT-BR pattern
  - New web/lib/i18n/pick.ts helper (unused yet by chat/agents — kept
    for future locale-driven switching when both languages are equally
    full; current rule is PT-BR-first since the user is brasileiro)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-05-24 12:02:59 -03:00

Luiz Gustavo

25f19aee63

W3.7 followup: harden Dupin scoping + chunk_id parsing

CI / Web — typecheck + lint + build (push) Failing after 32s

Details

CI / Scripts — Python smoke (push) Failing after 3s

Details

CI / Web — npm audit (push) Failing after 27s

Details

CI / Retrieval — golden set (Recall@5 + MRR) (push) Failing after 3s

Details

Two regressions surfaced in the smoke test that put Dupin from
0/3 contradictions written → 3/3 in the next run.

1. Single-doc scope was too narrow for Dupin's task.
   Holmes's question about Sandia returned 4 chunks scoped to one doc,
   but Dupin's terser "topic" form yielded only 1. Solution: Pass-1
   tries the requested doc_id; if the head is < 2 chunks, Pass-2
   widens to the whole corpus. Audit event carries `scope_widened`
   so the case-writer can later flag cross-doc contradictions
   distinctly. The unscoped retry hit 9 chunks and produced 3
   contradictions across 3 different docs.

2. Chunk-block header was ambiguous to the model.
   `--- doc-id/p007#c0042 ---` led Claude to parse `chunk_id` as
   "p007#c0042" or "p007/c0042" in the JSON output. write_contradiction
   then refused the FK lookup with "chunk not found". Fix:
   - Explicit `doc_id:` / `chunk_id:` / `page:` lines per chunk
     in the rendered block (no slashes/hashes the model can fold).
   - Defensive normalizeChunkId() in write_contradiction.ts strips
     any pNNN prefix and keeps only the trailing cNNNN — so the
     writer is forgiving without losing strictness on the topic +
     statement validation.

Smoke now produces (job 6deddf4b):
  R-0001 (3 chunks) — Color of the fireball(s) in incident summaries
  R-0002 (2 chunks) — Geographic confinement of green-fireball sightings
  R-0003 (3 chunks) — Whether the phenomenon was exclusively green or
                      also red/multicolored

R-0003 connects 3 different declassified documents: the Los Alamos
conference (exclusively-green category), a retrospective document
(red OR green), and Incident 229 (red, blue, yellow — no green).
Real cross-doc contradiction, fully cited.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-05-23 21:42:01 -03:00

Luiz Gustavo

5ac53cb3e2

W3.7: Dupin contradiction-scan detective + UI integration

CI / Web — typecheck + lint + build (push) Failing after 39s

Details

CI / Scripts — Python smoke (push) Failing after 4s

Details

CI / Web — npm audit (push) Failing after 37s

Details

CI / Retrieval — golden set (Recall@5 + MRR) (push) Failing after 4s

Details

Adds the third AI detective in the Investigation Bureau runtime: C. Auguste
Dupin, who scans a corpus shortlist for pairs (or small groups) of chunks
that cannot both be true under any ordinary reading.

Runtime:
  - prompts/dupin.md — discipline (no contradiction without ≥2 distinct
    chunk_ids; reject same-vocabulary near-misses; FEW high-confidence
    over MANY weak ones; emit `NO_CONTRADICTIONS` when corpus is silent)
  - src/detectives/dupin.ts — hybridSearch with k=18 (more chunks than
    Holmes because contradictions emerge from comparing dispersed
    claims), strict JSON-array parsing, AT MOST 3 contradictions per call
  - src/tools/write_contradiction.ts — validates topic + ≥2 positions
    drawn from ≥2 distinct chunks, resolves chunk_pk via DB lookup
    (rejects positions citing unknown chunks), INSERTs into
    public.contradictions + writes case/contradictions/R-NNNN.md
  - orchestrator: new `contradiction_scan` kind dispatching to runDupin;
    payload { topic, doc_id?, lang?, context_chunks? }

Chat + UI:
  - request_investigation gains kind=contradiction_scan + topic arg;
    triggered detective auto-resolves to dupin
  - chat-bubble inline card renders dupin in orange (#ff8a4d) to
    distinguish from holmes (cyan) and locard (green)
  - /jobs/[id] page swaps title + subtitle + tone per detective;
    "Question" label becomes "Topic" for contradiction_scan
  - /api/jobs/[id] hydrates public.contradictions when outputs[] surfaces
    contradiction_ids
  - job-status-poller renders ContradictionCard: topic + N positions
    (verbatim statements quoted, stance label optional, link to source
    chunk) + optional notes panel, with resolution_status badge
    (open/resolved/irreconcilable)

R-NNNN shares the contradiction_id_seq slot with relation per
CLAUDE.md naming — same conceptual class (a connection between two
pieces of evidence in tension).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-05-23 21:34:04 -03:00

3 commits