AI coding tools are now used by a majority of developers, and agentic use of these tools has popularized the practice colloquially called "vibe coding". Yet causal evidence on their effect on software architecture is scarce. Prior causal work has measured code-level outcomes (complexity, static analysis warnings); whether such degradation propagates to architecture-level outcomes remains unknown. We mine 151 open-source Java repositories, 74 with detectable agentic AI adoption (identified via configuration files and Co-Authored-By commit trailers) and 77 propensity-matched controls, across a 13-month per-repository window yielding 1,811 monthly Arcan snapshots. We estimate the causal effect of adoption on architectural smell density (ASD) with a staggered difference-in-differences design and the Borusyak imputation estimator, applying a causal design recently used for code-level metrics to the architecture level. Total smell counts are essentially unchanged (+1.1%, p = 0.82) while lines of code grow +12.8% (p = 0.003); the resulting 6.7% ASD decline (p = 0.004) is therefore a denominator effect rather than an architectural improvement. Per-type estimates and robustness checks (wild cluster bootstrap, Lee bounds, stale-observation sensitivity) corroborate the pattern; pre-trends are flat (Wald p = 0.90), consistent with parallel trends. Density-normalized outcomes can mislead when treatment affects system size: raw counts and explicit decomposition are required for causal mining studies of AI tool adoption. The complete replication package, including the curated 151-repository monthly panel, is publicly available.
翻译:人工智能编码工具如今已被大多数开发者使用,这些工具的智能体化应用也普及了俗称“氛围编码”的实践。然而,关于其对软件架构影响的因果证据仍然稀缺。以往的因果研究衡量了代码层面的结果(复杂度、静态分析警告);这种退化是否会传导至架构层面仍属未知。我们挖掘了151个开源Java仓库,其中74个可检测到采用了智能体化人工智能(通过配置文件及Co-Authored-By提交标记识别),并匹配了77个倾向性得分对照仓库,每个仓库跨越13个月的时间窗口,生成了1,811个月度Arcan快照。我们采用交错双重差分设计与Borusyak插值估计量,将近期用于代码级度量的因果设计应用于架构层面,估算了采用人工智能对架构气味密度(ASD)的因果效应。总气味计数基本不变(+1.1%,p = 0.82),而代码行数增长了+12.8%(p = 0.003);因此,由此产生的6.7%的ASD下降(p = 0.004)是分母效应而非架构改进。按类型的估计值与稳健性检验(聚束自举法、Lee边界、陈旧观测敏感性)均印证了这一模式;预处理趋势平坦(Wald p = 0.90),符合平行趋势假设。当处理措施影响系统规模时,密度归一化的结果可能产生误导:针对人工智能工具采用的因果挖掘研究需要原始计数与显式分解。完整的可复现包(包含精心整理的151仓库月度面板数据)已公开提供。