disclosure-bureau

Author	SHA1	Message	Date
Luiz Gustavo	a7e9dce6d2	rebuild entity layer from Sonnet-vision reextract pipeline Add reextract pipeline (scripts/reextract/) that rebuilds doc-level entity JSON from Sonnet-vision chunks via Opus, replacing the noisy per-page extraction. Add synthesize scripts to regenerate wiki/entities from the 116 _reextract.json (30), aggregate missing page.md from chunks (31), and reprocess 805 pages the doc-rebuilder agent dropped on context overflow (32). Add maintain scripts 43-56 for chunk-page sync, dedup, generic-entity marking, and typed relation extraction. Web: wire relations API + entity-relations component; entity/timeline/doc pages consume the rebuilt layer. Note: raw/, processing/, wiki/ remain gitignored (bulk data managed separately); the 116 reextract JSONs and 7,798 rebuilt entity files live on disk only. The 27 curated anchor events under wiki/entities/events/ are preserved. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-05-21 12:20:24 -03:00
guto	9889308bf4	fix chat: force synthesis pass + fix ambiguous-column trigger Two bugs combined to make the chat reply with only cards and no prose: 1. SQL trigger rollup_session_stats was failing with "column reference total_cost_usd is ambiguous" because the UPDATE on public.profiles had a FROM public.chat_sessions clause and both tables expose that column. Persistence of every user message died at this point — sessions were created in the DB but had message_count=0 forever. Applied SQL fix that qualifies columns with p./s. aliases (production DB updated; ALTER FUNCTION run live, not yet codified in a migration file). 2. The free-tier model (nemotron-3-super:free) spent all 5 tool-loop turns on hybrid_search calls and never wrote any prose, returning content_len=0. Added a forced-synthesis pass in openrouter.ts: when the loop exits with empty assembledText but the model did call tools, we send ONE final turn with tools omitted from the request payload and a user message instructing the model to answer in 3-8 sentences citing chunks. openrouterStreamCall now accepts a `withTools` opt so the synthesis call can disable tool calling entirely. Verified end-to-end with the actual user query "O que os astronautas viram? Quem foi que viu?" on /d/nasa-uap-d6-apollo-17-...: - content_len: 0 → 947 chars (real synthesis citing Schmitt) - artifacts: 44 preserved - assistant message persisted with tool_calls + citations columns	2026-05-18 15:39:46 -03:00
guto	d5f6e6030a	fix png-numbering: re-convert 34 zero-based docs + crop fallback 34 of 116 docs were generated with 0-based PNG numbering (p-000.png … p-008.png) but the Sonnet chunks reference 1-based page numbers in their YAML frontmatter (page: 9 means the 9th sheet of paper). The /api/crop handler built p-009.png and got a 500, the browser's Next/Image surfaced 400, and the chunk rendered as a black box on screen. Fixes: - web/app/api/crop/route.ts: try p-NNN.png first, fall back to p-(NNN-1).png if the 1-based file is missing. Cheap insurance for any doc that comes in with the old convention. - scripts/01-convert-pdfs.sh: previously printf '%03d' "$num" with $num starting at 0 (e.g. "008") raised "invalid number" because Bash parsed it as octal. Wrap with $((10#$num)) to force decimal — this was silently corrupting page sequences and producing gaps like p-001 ... p-008, p-011 (missing p-009/p-010). - All 34 affected docs re-converted from PDFs with the patched script; every directory now has continuous 1-based PNGs. - /processing/png/ rsync'd to VPS, web redeployed. Smoke: /api/crop?doc=doc-341-…&page=9&… now returns 200 image/webp instead of 500. Tested in browser: chunk c0026 (diagram, p9) renders the real engineering diagram.	2026-05-18 11:45:40 -03:00
guto	7d13f93393	ship: synthesize 158 entities, AG-UI artifacts, chat persistence, auth flow Fase 3 onda 2 — entity synthesis at scale: - scripts/synthesize/20_entity_summary.py: queries DB for entities with total_mentions ≥ threshold + top-K verbatim chunk snippets via entity_mentions JOIN, prompts Sonnet (Holmes-Watson voice, bilingual), writes narrative_summary EN+PT-BR + summary_status=synthesized. Ran on 187 candidates (mentions ≥ 20) → 158 OK · 1 err · 29 skipped (no snippets). Combined with anchor curation: 20 curated + 158 synthesized = 178 entities with real narrative (vs 0 a day ago). Fase 4 — chat with typed artifacts + persistence: - lib/chat/agui.ts: AG-UI v1 typed Artifact union (citation, crop_image, entity_card, evidence_card, hypothesis_card, case_card, navigation_offer) alongside the existing event types. - lib/chat/tools.ts + openrouter.ts: hybrid_search emits up to 6 citation + crop_image artifacts per query. Provider collects them and returns in done.artifacts so the route can persist. - api/sessions/[id]/messages: persist artifacts to messages.citations. - components/chat-bubble.tsx: ArtifactCard renders inline cards (citation, crop_image, entity_card, navigation_offer) for streamed and persisted messages. activeId now persisted in localStorage so navigation between pages keeps the same conversation. New sessions are lazy (only when user has zero). loadMessages hydrates tools + artifacts from server. CRUD UI: rename (✎) + archive (🗑) buttons per session in the list. Home search: - doc-list-filters: input now fires hybrid_search (rerank=0 for speed) in parallel with the local title filter; chunk hits render above the doc grid with snippet + score + classification. - api/search/hybrid: accept ?rerank=0 to skip the cross-encoder (1.3s vs 60s). Auth flow: - infra: SMTP_HOST=mail.spacemail.com:587 + DMARC published; mail now lands in inbox. GOTRUE_MAILER_AUTOCONFIRM=false (real email verification). - kong.yml: proxy /auth/callback on api.disclosure.top → web:3000 so PKCE email links don't 404 at the gateway. - web/app/auth/callback: handle both ?code= (OAuth) and ?token=&type= (PKCE); redirect to the public site host before verifyOtp so the session cookie lands on the right domain. Audit deliverables: - .nirvana/outputs/disclosure-bureau/.../systems-atelier/: 5 docs (code analysis, tech debt, discovery brief, system arch, 5 ADRs) authored by sa-principal that produced this roadmap. Kept in-tree for traceability.	2026-05-18 03:52:59 -03:00
guto	4459bd17e4	phase-0: kill stubs, ship 20 curated anchor events, configure SMTP - scripts/03-dedup-entities.py: stop emitting placeholder narrative ("Stub. Will be enriched in Phase 7"); write summary_status=none + null fields instead. - scripts/maintain/41_strip_stubs.py: idempotent migration that cleaned the 22,096 entity .md files (now zero stub strings in wiki/). - scripts/synthesize/01_anchor_events.py: curated 20 anchor UAP events (Roswell, Nimitz Tic-Tac, Phoenix Lights, Operação Prato, AATIP, etc.) with bilingual Holmes-Watson narrative via claude -p --model sonnet (CLAUDE_CODE_OAUTH_TOKEN). All summary_status=curated, confidence=high. - web/api/timeline + timeline-view: filter narrative-less events by default, render "curado" badge for hand-vetted ones, drop the date display alone. - CLAUDE-schema-full.md: document the summary_status enum and the four states. - docker-compose.yml: SMTP_HOST=mail.spacemail.com configured; GOTRUE_MAILER_AUTOCONFIRM flipped to false (real email confirmation working). - .nirvana/outputs/.../systems-atelier/: 5 deliverables of the architecture audit that produced this roadmap.	2026-05-18 00:44:17 -03:00
guto	19d0678e55	baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack	2026-05-17 22:44:36 -03:00

6 commits