成为VIP会员查看完整内容
VIP会员码认证
首页
主题
会员
服务
注册
·
登录
Claude
关注
2
综合
百科
VIP
热门
动态
论文
精华
Invalidation Contracts for Cross-Episode Agent Memory
Arxiv
0+阅读 · 8月31日
SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task
Arxiv
0+阅读 · 9月1日
MemoryWalker: Stop Training Agents on Contexts They Never Saw
Arxiv
0+阅读 · 9月1日
On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces
Arxiv
0+阅读 · 8月28日
An Empirical Evaluation of Using Large Language Models for Automated Model-Based Test Generation
Arxiv
0+阅读 · 8月27日
FABRICA: Agentic CUDA-to-CSL Translation and Optimization for Wafer-Scale Systems
Arxiv
0+阅读 · 8月25日
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
Arxiv
0+阅读 · 8月12日
SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification
Arxiv
0+阅读 · 7月17日
Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI
Arxiv
0+阅读 · 7月1日
When Large Language Models are More PersuasiveThan Incentivized Humans, and Why
Arxiv
0+阅读 · 8月13日
Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa?
Arxiv
0+阅读 · 7月22日
AI Security Leaderboard: Methodology, Results and Minimal Standard
Arxiv
0+阅读 · 8月4日
Code Semantic Zooming
Arxiv
0+阅读 · 6月30日
Streaming Communication in Multi-Agent Reasoning
Arxiv
0+阅读 · 8月1日
FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs
Arxiv
0+阅读 · 8月25日
参考链接
提示
微信扫码
咨询专知VIP会员与技术项目合作
(加微信请备注: "专知")
微信扫码咨询专知VIP会员
Top