The exponential growth of scientific literature, datasets, and code repositories has created a discovery bottleneck that impedes knowledge synthesis and reproducibility. Traditional dissemination formats -- static PDFs, siloed code hosting, and fragmented data repositories -- fail to represent the interconnected narrative of modern research, while conventional metrics such as the H-index neglect contributions from reusable code and shared datasets. We present ResearchTwin, an open-source federated platform that transforms a researcher's scholarly output into a conversational digital twin, with a preliminary evaluation of its deployed prototype. The system uses a Bimodal Glial-Neural Optimization (BGNO) architecture comprising a Multi-Modal Connector Layer, a Glial Layer for caching and rate management, and a Neural Layer implementing Retrieval-Augmented Generation with a provider-agnostic LLM backend. We formalize the S-index, building on our earlier QIC framework, into a composite metric that extends FAIR principles -- via a binary accessibility/licensing gate, field-normalized impact scoring, and geometric collaboration scaling -- to quantify multimodal research impact. A case study comparing two researchers with similar H-indexes but substantially different S-indexes demonstrates that the metric captures dimensions of impact -- particularly dataset and code contributions -- invisible to citation-based measures alone. ResearchTwin exposes an inter-agentic discovery API using Schema.org typed responses and HATEOAS navigation, enabling AI agents to discover cross-lab synergies. A three-tier federated architecture preserves data sovereignty while enabling global discoverability.
翻译:科学文献、数据集与代码仓库的指数级增长造成了阻碍知识综合与可复现性的发现瓶颈。传统的传播形式(静态PDF、孤立的代码托管平台及碎片化数据仓库)无法体现现代研究的互连叙事特征,而诸如H指数等传统指标则忽视了可复用代码与共享数据集的学术贡献。本文提出ResearchTwin——一个将研究人员学术产出转化为可对话数字孪生的开源联邦平台,并对其实例化原型进行了初步评估。该系统采用双模态胶质-神经优化(BGNO)架构,包含多模态连接层、用于缓存与速率管理的胶质层,以及实现基于检索增强生成(RAG)的神经层(配备中立型大语言模型后端)。我们在前期QIC框架基础上,将S指数形式化为一个复合指标——通过二进制可访问性/许可门控、领域归一化影响力评分及几何协作缩放扩展了FAIR原则——以量化多模态研究影响力。对两位H指数相近但S指数差异显著的研究人员的案例研究表明,该指标能够捕捉纯引用指标无法体现的影响力维度(特别是数据集与代码贡献)。ResearchTwin通过Schema.org类型化响应与HATEOAS导航暴露了智能体间发现API,使AI智能体能够发现跨实验室协同效应。其三层级联邦架构在保障数据主权的同时实现了全局可发现性。