成为VIP会员查看完整内容
VIP会员码认证
首页
主题
会员
服务
注册
·
登录
声明
关注
0
综合
百科
VIP
热门
动态
论文
精华
How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation
Arxiv
0+阅读 · 9月1日
Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free
Arxiv
0+阅读 · 8月29日
Frontiers in FinTech: Multimodal Foundation Models for Financial Reporting and Decision Science
Arxiv
0+阅读 · 8月30日
Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair
Arxiv
0+阅读 · 9月1日
LOOMSUM:Weaving Quantitative and Narrative Evidence for Faithful Long Text-Table Summarization
Arxiv
0+阅读 · 8月31日
Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection
Arxiv
0+阅读 · 8月30日
Ontology-Guided Multi-Agent Extraction of Evaluation Objects from Academic Review Texts: Evidence from Chinese Library and Information Science
Arxiv
0+阅读 · 8月30日
Beyond Good Intentions: When Does the Framing of Multilingual and Low-Resource NLP Research Become a Caricature?
Arxiv
0+阅读 · 8月31日
Entropy lower bounds and sum-product phenomena
Arxiv
0+阅读 · 8月28日
Fine Difference Structure and Prime-Power Depth of Bent Partitions
Arxiv
0+阅读 · 8月28日
A Constant Metric Distortion Protocol for Approval Voting Given Plurality Polls
Arxiv
0+阅读 · 8月28日
Mean-Tilted Intervals: Short Tolerance Intervals
Arxiv
0+阅读 · 8月28日
It Takes Three to Converse: Empirical Observations on How the Developer, the Convener and the Participant Shaped 119 Polis Conversations
Arxiv
0+阅读 · 8月28日
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
Arxiv
0+阅读 · 8月31日
SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization
Arxiv
0+阅读 · 8月29日
参考链接
提示
微信扫码
咨询专知VIP会员与技术项目合作
(加微信请备注: "专知")
微信扫码咨询专知VIP会员
Top