Research methodology and sources
Web search is the start, not the method. Orient, go to authoritative sources, chain citations, corroborate across independent sources, and record provenance as you go.
Workflow
- Orient and classify the topic's domain.
- Identify authoritative sources for that domain (catalogue below).
- Search with Boolean or controlled vocabulary (e.g. MeSH for medicine).
- Chain citations — backward and forward snowballing — to saturation.
- Corroborate across genuinely independent sources.
- Record provenance at generation time. Use PRISMA when systematic rigor is required.
Source catalogue
- Academic: OpenAlex (free, no key — a strong default), Semantic Scholar, PubMed / Entrez, arXiv, Crossref, OpenCitations, CORE, Unpaywall, Scite.
- Primary / official: RFC/IETF, W3C, NIST/NVD, USPTO, SEC EDGAR, data.gov, World Bank, Eurostat, Census, ClinicalTrials.gov, CourtListener/RECAP, Congress.gov.
- Domain / community: official docs + GitHub + Stack Exchange + package registries (technical); Cochrane / medRxiv (medical); FRED / BLS (finance); Chronicling America / DPLA / Europeana (history).
Agentic access (APIs and MCP)
Most academic and government sources expose free REST APIs an agent can call directly, and many have ready-made MCP servers (arxiv, semantic-scholar, openalex, pubmed, github). Agentic search layers include Tavily, Exa, Firecrawl, and Perplexity Sonar. Prefer reading full sources over snippets.