The Disclosure Bureau — declassified UAP/UFO wiki
Find a file
guto 7d13f93393 ship: synthesize 158 entities, AG-UI artifacts, chat persistence, auth flow
Fase 3 onda 2 — entity synthesis at scale:
- scripts/synthesize/20_entity_summary.py: queries DB for entities with
  total_mentions ≥ threshold + top-K verbatim chunk snippets via
  entity_mentions JOIN, prompts Sonnet (Holmes-Watson voice, bilingual),
  writes narrative_summary EN+PT-BR + summary_status=synthesized.
  Ran on 187 candidates (mentions ≥ 20) → 158 OK · 1 err · 29 skipped (no
  snippets). Combined with anchor curation: 20 curated + 158 synthesized
  = 178 entities with real narrative (vs 0 a day ago).

Fase 4 — chat with typed artifacts + persistence:
- lib/chat/agui.ts: AG-UI v1 typed Artifact union (citation, crop_image,
  entity_card, evidence_card, hypothesis_card, case_card, navigation_offer)
  alongside the existing event types.
- lib/chat/tools.ts + openrouter.ts: hybrid_search emits up to 6
  citation + crop_image artifacts per query. Provider collects them and
  returns in done.artifacts so the route can persist.
- api/sessions/[id]/messages: persist artifacts to messages.citations.
- components/chat-bubble.tsx: ArtifactCard renders inline cards (citation,
  crop_image, entity_card, navigation_offer) for streamed and persisted
  messages. activeId now persisted in localStorage so navigation between
  pages keeps the same conversation. New sessions are lazy (only when user
  has zero). loadMessages hydrates tools + artifacts from server. CRUD UI:
  rename (✎) + archive (🗑) buttons per session in the list.

Home search:
- doc-list-filters: input now fires hybrid_search (rerank=0 for speed)
  in parallel with the local title filter; chunk hits render above the doc
  grid with snippet + score + classification.
- api/search/hybrid: accept ?rerank=0 to skip the cross-encoder (1.3s vs 60s).

Auth flow:
- infra: SMTP_HOST=mail.spacemail.com:587 + DMARC published; mail now lands
  in inbox. GOTRUE_MAILER_AUTOCONFIRM=false (real email verification).
- kong.yml: proxy /auth/callback on api.disclosure.top → web:3000 so PKCE
  email links don't 404 at the gateway.
- web/app/auth/callback: handle both ?code= (OAuth) and ?token=&type=
  (PKCE); redirect to the public site host before verifyOtp so the session
  cookie lands on the right domain.

Audit deliverables:
- .nirvana/outputs/disclosure-bureau/.../systems-atelier/: 5 docs (code
  analysis, tech debt, discovery brief, system arch, 5 ADRs) authored by
  sa-principal that produced this roadmap. Kept in-tree for traceability.
2026-05-18 03:52:59 -03:00
.nirvana/outputs/disclosure-bureau phase-0: kill stubs, ship 20 curated anchor events, configure SMTP 2026-05-18 00:44:17 -03:00
infra ship: synthesize 158 entities, AG-UI artifacts, chat persistence, auth flow 2026-05-18 03:52:59 -03:00
scripts ship: synthesize 158 entities, AG-UI artifacts, chat persistence, auth flow 2026-05-18 03:52:59 -03:00
web ship: synthesize 158 entities, AG-UI artifacts, chat persistence, auth flow 2026-05-18 03:52:59 -03:00
wiki/entities/events phase-0: kill stubs, ship 20 curated anchor events, configure SMTP 2026-05-18 00:44:17 -03:00
.gitignore baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack 2026-05-17 22:44:36 -03:00
CLAUDE-schema-full.md phase-0: kill stubs, ship 20 curated anchor events, configure SMTP 2026-05-18 00:44:17 -03:00
CLAUDE.md baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack 2026-05-17 22:44:36 -03:00
CORPUS-SNAPSHOT.md baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack 2026-05-17 22:44:36 -03:00
README.md baseline: Disclosure Bureau pipeline + Next.js UI + Supabase stack 2026-05-17 22:44:36 -03:00

The Disclosure Bureau

Investigative wiki + agentic chat sobre o corpus declassificado do US Department of War em war.gov/ufo (116 PDFs, 3.435 páginas, 34k+ entidades, 28 vídeos UAP).

Live: disclosure.top

O que é

Pipeline de IA que transforma documentos UAP/UFO declassificados em uma wiki investigativa navegável + chat agêntico com retrieval semântico bilíngue (EN + PT-BR) e citações com bbox crop no PDF original.

A premissa metodológica é o padrão Karpathy LLM Wiki: ler tudo, compilar conhecimento em markdown cross-referenciado, navegar via wiki-links — não por busca vetorial. Em cima dessa wiki rodamos uma camada de hybrid retrieval (BM25 + BGE-M3 dense + cross-encoder rerank) para perguntas livres no chat.

A camada investigativa segue protocolo Investigation Bureau (Holmes/Poirot/Dupin/Locard + Schneier/Tetlock/Taleb): chain-of-custody, hypothesis tournament, residual uncertainty.

Arquitetura

PDFs (raw/)
  ↓ pdftoppm 72 DPI + pdftotext
processing/ (png + ocr)
  ↓ Sonnet 4.6 subagents (page-rebuilder, image-analyst, table-stitcher)
raw/<doc>--subagent/ (chunks bilíngues + bbox + anomaly flags)
  ↓ scripts/30 (BGE-M3 embed) + 31 (entity_mentions)
Postgres + pgvector + tsvector
  ↓ hybrid_search RPC + reranker
chat agente (OpenRouter) cita [[doc/p007#c0042]] → frontend renderiza crop bbox

Stack

  • Embedding: BGE-M3 self-hosted (1024-dim, multilíngue, $0)
  • Reranker: BGE-Reranker-v2-M3 self-hosted ($0)
  • Vetor + texto: Postgres 15 + pgvector + tsvector bilíngue (pt_unaccent, en_unaccent)
  • LLM (chat): OpenRouter — DeepSeek v4 free como default
  • Frontend: Next.js 15 + React 19 + Tailwind + assistant-ui (Pattern C streaming)
  • Auth + persistência: Supabase self-hosted (GoTrue, PostgREST, Storage, Imgproxy)
  • Reverse proxy: Traefik + Let's Encrypt
  • Imagens: sharp via /api/crop (bbox on-demand, cached 1ano)

Layout

/Users/guto/ufo/
├── CLAUDE.md                       # contrato vinculante (24 tipos de markdown)
├── CLAUDE-schema-full.md           # schema detalhado
├── README.md                       # este arquivo
├── raw/                            # 116 PDFs imutáveis + chunks v0.2.0 derivados
│   ├── <pdfs>
│   ├── <doc-id>--subagent/         # chunks rebuilt (chunks/c*.md + _index.json + document.md)
│   └── _batch-rebuild/             # logs do orchestrator
├── processing/                     # intermediários (PNG, OCR, vision JSON)
├── wiki/                           # markdown gerado (documents/, pages/, entities/, tables/, images/)
├── case/                           # artefatos Investigation Bureau (case-report, hypotheses, gaps)
├── scripts/                        # 33 scripts numerados (Phase 0 → manutenção)
├── infra/                          # docker-compose, embed-service, migrations, deploy
└── web/                            # Next.js frontend

Quick start

# 1. Converter PDFs em PNG + OCR (uma vez)
./scripts/01-convert-pdfs.sh

# 2. Rebuild chunks bilíngues (Sonnet 4.6 via Claude Code subagents)
python3 scripts/28-batch-rebuild-all.py --workers 2

# 3. (após batch) Indexar em Postgres + embeddings
python3 scripts/30-index-chunks-to-db.py --skip-existing

# 4. (opcional) Materializar entity_mentions p/ grafo
python3 scripts/31-populate-entity-mentions.py

# 5. Deploy
cd infra/disclosure-stack && ./scripts/deploy.sh

Detalhes completos em infra/DEPLOY-CHECKLIST.md.

Features do frontend

URL Função
/ Lista de documentos com resumo de 3 linhas, filtros (collection, classification, sort), busca
/d/<doc> Visão legado (page grid + frontmatter)
/d/<doc>/v2 Render rico de chunks com lang toggle (PT/EN/both), paged vs flow
/d/<doc>/v2/<page> Single page V2 com PNG side-by-side
/d/<doc>/full Texto consolidado bilíngue
/e/<class> Lista paginada de entidades por classe (people, locations, ...)
/e/<class>/<id> Detalhe da entidade + co-mentions + chunks live
/search?q=... Hybrid search URL-shareable
/timeline Cronologia de eventos por década
/graph Grafo força-direcionado de co-menções (Obsidian-style)
/admin/stats Analytics do corpus (FS + DB)
/admin/batch Monitor de progresso do rebuild
/admin/indexer Estado da camada de retrieval

Atalhos globais:

  • ⌘K / Ctrl+K em qualquer página → command palette com hybrid_search
  • Toggle 🌐 EN ↔ PT-BR fixo bottom-left (cookie 1ano)
  • Chat 💬 botão flutuante bottom-right com 12 ferramentas

Os 12 tools do agente

🔍 Retrieval: hybrid_search, read_chunk, get_page_chunks, list_anomalies 🔗 Grafo: entity_neighbors, entity_path, co_mention_chunks 📄 Wiki: read_document, read_page, read_entity, search_corpus 🧭 UI: navigate_to

Citações tipo [[doc-id/p007#c0042]] viram cards interativos com crop bbox + texto bilíngue + link.

Custos

Item Custo
Rebuild chunks (Sonnet 4.6 via Claude Code Max 20x) ~$200 one-shot p/ 116 docs
Embedding BGE-M3 self-host $0/mês
Reranker BGE-Reranker-v2-M3 self-host $0/mês
Postgres + pgvector já incluso no VPS
Chat LLM (DeepSeek free via OpenRouter) $0/req
VPS (16GB / 4 CPU) ~€10/mês

Documentação

Licença + procedência

  • PDFs declassificados: domínio público (US Department of War / FBI / DOS / NASA)
  • Código deste projeto: MIT
  • Modelos: BGE-M3 (MIT), DeepSeek v4 (proprietary via OpenRouter free tier)
  • Branding: The Disclosure Bureau / disclosure.top — pessoal

Wiki investigativa, não advocacy. Toda claim tem chain-of-custody até a página + bbox do PDF original.