disclosure-bureau

History

Luiz Gustavo e75ca5eda2 add clean LLM reading version of documents (the core goal) Scanned docs are messy — duplicate transcriptions (typed + handwritten), two classification variants of the same narrative, OCR noise, repeated banners. The doc page showed raw chunks, so everything appeared twice. 40_reading_version.py generates ONE clean, deduplicated, well-structured bilingual Markdown reading version per doc (Sonnet): merges duplicate versions without losing unique lines, drops page furniture, formats transcripts as dialogue. Faithful — invents nothing; redactions kept as markers. /d/[docId] now defaults to a "📖 leitura" tab rendering this clean version, with "🔍 trechos · scan original" preserving the faithful per-chunk + per-page scan view. reading.md lives in raw/<doc>--subagent/ alongside the chunks. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>		2026-05-21 17:23:36 -03:00
..
chat	guard /admin/* by role + filter chat artifacts to cited chunks	2026-05-18 17:41:35 -03:00
retrieval	search: gate dense recall by cosine-distance threshold in the RPC	2026-05-21 16:36:56 -03:00
supabase	baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack	2026-05-17 22:44:36 -03:00
chunks.ts	add clean LLM reading version of documents (the core goal)	2026-05-21 17:23:36 -03:00
doc-renderer.ts	baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack	2026-05-17 22:44:36 -03:00
doc-summary.ts	baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack	2026-05-17 22:44:36 -03:00
entity-index.ts	baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack	2026-05-17 22:44:36 -03:00
fm-types.ts	baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack	2026-05-17 22:44:36 -03:00
wiki.ts	baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack	2026-05-17 22:44:36 -03:00