NTH

How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming

AuthorsGabrielle O'Brien, Reed Milewicz, Nasir Eisty

AffiliationsSandia National Laboratories, Albuquerque, New Mexico, USA · University of Tennessee, Knoxville, Knoxville, Tennessee, USA

September 22, 2026 3 min read
Watch on YouTube
The one-line take

A survey of 527 researchers finds that AI-generated scientific code is usually checked informally by individuals rather than through systematic tests or peer review.

Key results

527
Analyzed AI-use accounts

Self-reported episodes of generative AI use in scientific programming.

76%
Top-five task share

Share of coded use cases involving data handling, visualization, debugging, mathematical or scientific computing, and statistical analysis.

52.8%
Ran generated code

Accounts reporting manual execution as an evaluation strategy.

2.8%
Automated test suites

Accounts explicitly mentioning unit, integration, or other test-harness checks.

2.0%
Colleague review

Accounts reporting review by another person.

3.7
Confidence crossover

Approximate years of programming experience where confidence shifted from favoring the AI to favoring the programmer.

What the paper found

This study analyzes 527 self-reported episodes in which scientific programmers used generative AI for research code. The dominant tasks were data handling, visualization, debugging, mathematical or scientific computing, and statistical analysis, together representing 76% of coded use cases. Respondents most often named ChatGPT, followed by GitHub Copilot, Google Gemini, and Claude, while Claude Code assisted the study’s quantitative analysis scripts. Validation was usually an informal, closed-loop process: 52.8% of accounts said they ran the generated code, often checking for errors or whether outputs looked plausible, while reading code, inspecting plots, consulting documentation, and comparing against benchmarks were less common. Independent safeguards were rare: automated unit or integration test suites appeared in 2.8% of accounts and review by a colleague in 2.0%. Programming experience had little relationship with the tasks delegated or evaluation strategies reported, but it strongly shaped confidence. Less experienced programmers tended to trust the AI more than themselves, whereas experienced programmers showed the opposite pattern; the confidence crossover occurred at approximately 3.7 years of programming experience. Confidence in evaluating results was not reliably associated with using more validation strategies, and instead correlated most strongly with confidence in the tool and in one’s own ability. The paper argues that AI coding interfaces should generate task-specific verification artifacts, such as regression cases for debugging, diagnostic views for visualization, and known-answer comparisons for scientific computation, rather than leaving correctness judgments entirely to individual researchers.

Original abstract

Generative AI has entered research programming, yet there is little evidence about which tasks researchers hand to it or how they decide whether its code is correct. We draw on 527 free-text responses to a 2025 survey of researchers who write code, most of them at U.S. universities. In each response, a researcher recounts a single task from their own work, the way they used an AI tool for it, and what they did to assess the result. We coded the task and the evaluation strategies reported, and related both to programming experience, research area, and confidence ratings. Use was concentrated in five tasks: data handling, visualization, debugging, mathematical/scientific computing, and statistical analysis. Evaluation was informal and individual. Over half of accounts described running the generated code, while automated tests and review by another person were rare. Use cases and evaluation strategies varied little with programming experience, but confidence did: Less experienced programmers trusted the AI more than themselves, and experienced programmers the reverse. Evaluation confidence was not associated with the strategies reported. Its strongest correlates were confidence in the tool and in oneself. Validating AI contributions to scientific code rested largely on individual judgment, outside shared infrastructure for testing or review. Interfaces could support task-appropriate evaluation rather than leave it to the user.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis