SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
| Source: arXiv AI
Tags: scientific discovery, local LLMs, research automation, knowledge graphs, open-weight models, AI scientists
SGHA discovers scientific research gaps using only a local 9B open-weight model and a structured evidence graph — no proprietary API calls — producing traceable research problems with assumptions, objectives, and success criteria, compared favorably against AI Scientist-v2 across 5 machine-learning domains.
Details
Fully automated AI scientists typically rely on frontier models from proprietary APIs, creating data-governance risks when confidential research materials are transmitted externally. SGHA (Structural Gap Hypothesis Agent) bypasses this by running entirely on a locally served 9B open-weight model. The system structures a literature corpus into evidence-linked paper objects and a typed evidence graph, then detects unresolved structural patterns across papers before formulating research problems. Each output includes assumptions, objectives, success criteria, and remaining ambiguities — making the reasoning auditable rather than opaque. Candidate gaps are screened before formulation, reducing hallucination by grounding generation in explicit corpus structure. The authors compare SGHA against the AI Scientist-v2 idea formulation module across five machine-learning subdomains. Results suggest SGHA can match frontier-model quality on research problem formulation while keeping data on-premises. The paper acknowledges these are preliminary comparisons without statistical significance testing. For enterprise R&D teams, pharma researchers, or institutions with proprietary data that cannot be shared with external APIs, SGHA demonstrates a path to AI-assisted hypothesis discovery that respects data boundaries. The corpus-first, evidence-constrained design also produces outputs that are easier to audit for scientific validity.