NASA Ames Researchers Propose LLM-Based Decision-Support System for Aviation, Achieving Up to 100% Tool Selection Accuracy
Why It MattersThe research signals a cautious, tool-grounded path for integrating generative AI into aviation operations as advanced air mobility traffic grows more complex.
Researchers at NASA Ames Research Center have proposed a Large Language Model-based system to support decision-making in aviation operations, aimed particularly at managing complex airspaces as advanced air mobility technologies such as air taxis and delivery drones become more prevalent. Authored by Aida Sharif Rohani, Stephen S. Clarke, and Aditya Das, the system interprets natural language queries from users and routes them to validated aviation expert tools, including models for trajectory prediction, runway configuration, and Traffic Management Initiative forecasting, drawing on real-time data sources. Grounding LLM outputs in domain-validated tools rather than general-purpose language generation is intended to reduce the risk of hallucination.

The system was evaluated on a dataset of 1,000 manually drafted and verified aviation queries spanning nine tools and five question categories, including edge cases and deliberate wrong-tool prompts. Tests using GPT-4o and GPT-4o-mini produced tool selection accuracy ranging from 86% to 100% across all nine tools, validated on a held-out test set.
The paper, received 11 February 2026 and published 14 September 2026, notes that current FAA guidelines call for incremental AI deployment with robust output verification before operational use. The researchers found that aviation queries from professionals are often underspecified, using domain shorthand that omits required parameters, requiring the system to handle multi-turn clarification rather than assuming defaults.

















































