Scientific production in the era of large language models: Outcome-triggered treatment timing and spurious event-study dynamics.
Journal:
Proceedings of the National Academy of Sciences of the United States of America
Published Date:
Aug 11, 2026
Abstract
Large language models (LLMs) are increasingly used in scientific writing, but their effect on individual productivity is difficult to identify because adoption is rarely directly observed. [K. Kusumegi et al., Science 390, 1240-1243 (2025)] infer adoption from the first paper detected as LLM-assisted and report large productivity gains after adoption. We show that this treatment-timing rule mechanically generates positive event-study dynamics even in the absence of any causal effect. Because high-output months are more likely to produce a detected paper, treatment assignment becomes intrinsically linked to productivity. Using reconstructed arXiv data, we show that random treatment assignments, neutral-keyword triggers, inverted treatment, and pre-ChatGPT placebo periods all generate similar dynamics. Simulations with no treatment effect also reproduce the same posttreatment patterns reported in K. Kusumegi et al., Science 390, 1240-1243 (2025). Our results demonstrate that first-detection timing alone creates spurious evidence of productivity gains from LLM adoption.
Authors
Keywords
No keywords available for this article.