Legacy systems concentrate business rules, architectural decisions, and operational exceptions that often remain implicit in code, data, configuration, and maintenance practices. At the same time, language-model-based coding agents depend on reliable context, correctness criteria, and behavioral contracts to modify real systems with lower risk. This paper presents Reversa, a reverse documentation engineering framework for converting legacy software into traceable operational specifications for AI agents. Reversa organizes this process as a multi-agent pipeline: specialized agents map the project surface, analyze modules, extract implicit rules, synthesize architecture, write unit-level specifications, and review generated claims. The proposal emphasizes three mechanisms: traceability between code and specification, explicit confidence marking, and preservation of gaps for human validation. The framework is distributed as a Node.js CLI, installs skills across multiple agent engines, and uses a SHA-256 manifest to preserve modified files during update or uninstall operations. In addition to the architectural description, we report an exploratory case study on migrating an ATM from COBOL to Go, in which the pipeline produced 517 claims classified by an internal confidence index, 10 registered gaps, 53 Gherkin parity scenarios, and a reconstruction plan with 9 of 11 tasks completed at inventory time. Final parity validation and cutover were not completed in this study. We do not claim broad empirical superiority; we position the contribution with respect to the literature on reverse engineering, LLM-based documentation, and software agents, and propose an evaluation protocol with metrics for coverage, traceability, confidence, utility, and cost.
翻译:遗留系统集中了业务规则、架构决策和运行异常,这些内容通常隐含在代码、数据、配置和维护实践中。与此同时,基于语言模型的编码智能体依赖于可靠的上下文、正确性标准和行为契约,以较低风险修改真实系统。本文提出了Reversa,一个用于将遗留软件转化为AI智能体可追踪操作规范的反向文档工程框架。Reversa将此过程组织为多智能体流水线:专业智能体负责映射项目表面、分析模块、提取隐式规则、综合架构、编写单元级规范并审查生成的声明。该方案强调三种机制:代码与规范之间的可追溯性、显式置信度标记以及为人工验证保留的空白点。该框架以Node.js CLI形式分发,可在多个智能体引擎中安装技能,并使用SHA-256清单在更新或卸载操作期间保留已修改文件。除架构说明外,我们报告了一个将ATM从COBOL迁移至Go的探索性案例研究,其中流水线生成了517条按内部置信度索引分类的声明、10个已注册空白点、53个Gherkin并行场景,以及一份在盘点时已完成11项任务中9项的重建计划。本研究未完成最终的并行验证和切换。我们不声称具有广泛的实证优越性;我们将该贡献定位在与逆向工程、基于LLM的文档以及软件智能体相关的文献之中,并提出了一个包含覆盖度、可追溯性、置信度、效用与成本指标的评估协议。