ECAS: An Edge-Controlled Agentic System for Validation-Gated Scientific Application Execution

| Source: arXiv AI

Tags: LLM agents, HPC, scientific computing, edge computing, agentic systems, ALCF

ECAS separates LLM reasoning from HPC execution using an edge agent that retains credentials and validates artifacts before running at scale — improving scientific computing success from 0/6 to 6/6 over one-shot generation on two production ALCF systems.

Details

LLM agents can automate scientific computing workflows, but two obstacles block deployment: giving cloud-hosted models direct HPC access exposes credentials and execution authority; and one-shot code generation fails when artifacts encounter site-specific environment issues at scale.\n\nECAS (Edge-Controlled Agentic System) separates three concerns: a cloud LLM proposes plans and repairs; a user-controlled edge agent holds credentials, enforces policy, and gates execution; the HPC cluster computes. Generated artifacts pass static checks and small-scale validation before target-scale runs are permitted — failures trigger repair loops from sanitized feedback.\n\nExperiments on three scientific applications across two production ALCF systems (Argonne Leadership Computing Facility) with six injected fault types show: closed-loop repair improves success from 0/6 to 6/6 versus one-shot generation; validation gating prevents all three observed large-scale failures; edge-resident skill conditioning improves success from 4/6 to 6/6.\n\nThe architecture is directly applicable to enterprise HPC environments where security requirements prevent cloud model access to production systems.