Automated benchmarks dominate the evaluation of large language models, yet no systematic study has compared user satisfaction, adoption motivations, and frustrations across competing platforms using a consistent instrument. We address this gap with a cross-platform survey of 388 active AI chat users, comparing satisfaction, adoption drivers, use case performance, and qualitative frustrations across seven major platforms: ChatGPT, Claude, Gemini, DeepSeek, Grok, Mistral, and Llama. Three broad findings emerge. First, the top three platforms (Claude, ChatGPT, and DeepSeek) receive statistically indistinguishable satisfaction ratings despite vast differences in funding, team size, and benchmark performance. Second, users treat these tools as interchangeable utilities rather than sticky ecosystems: over 80% use two or more platforms, and switching costs are negligible. Third, each platform attracts users for different reasons: ChatGPT for its interface, Claude for answer quality, DeepSeek through word-of-mouth, and Grok for its content policy, suggesting that specialization, not generalist dominance, sustains competition. Hallucination and content filtering remain the most common frustrations across all platforms. These findings offer an early empirical baseline for a market that benchmarks alone cannot characterize, and point toward competitive plurality rather than winner-take-all consolidation among engaged users.


翻译:自动化基准测试主导了大型语言模型的评估,但尚无系统性研究采用一致性工具比较用户在竞争平台间的满意度、采用动机与挫败感。我们通过一项涵盖388名活跃AI聊天用户的跨平台调查弥补这一空白,比较了七个主要平台(ChatGPT、Claude、Gemini、DeepSeek、Grok、Mistral和Llama)的满意度、采用驱动因素、用例表现及定性挫败感。三大核心发现浮出水面:第一,尽管资金规模、团队规模和基准表现存在巨大差异,位列前三的平台(Claude、ChatGPT和DeepSeek)获得的满意度评分在统计上无显著差异;第二,用户将这些工具视为可互换的实用工具而非黏性生态:超80%用户同时使用两个及以上平台,且切换成本可忽略不计;第三,各平台吸引用户的动因各异:ChatGPT凭借界面设计,Claude依靠回答质量,DeepSeek通过口碑传播,而Grok因内容政策获客——这表明专业化竞争而非通用型主导方能维持市场活力。幻觉现象与内容过滤仍是所有平台最常见的用户挫败因素。这些发现为仅凭基准测试无法表征的市场提供了早期实证基线,并指向活跃用户群体中呈现竞争多元化而非赢家通吃的整合趋势。

0
下载
关闭预览

相关内容

【ChatGPT系列报告】ChatGPT/GPT-4 如何赋能应用,31页pdf
专知会员服务
167+阅读 · 2023年4月9日
AIGC行业深度报告:ChatGPT:重新定义搜索“入口”
专知会员服务
139+阅读 · 2023年2月10日
【Facebook】人工智能基准(Benchmarking)测试再思考,55页ppt
专知会员服务
32+阅读 · 2020年12月20日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
相关主题
最新内容
《决策模型比较研究》
专知会员服务
8+阅读 · 今天5:16
《美军水下战与海床战概述及本地实施》
专知会员服务
4+阅读 · 今天4:30
面向未来冲突推进陆军情报体制改革
专知会员服务
4+阅读 · 今天4:12
乌克兰纵深打击如何重塑俄罗斯的战略选择
专知会员服务
3+阅读 · 7月24日
俄乌战争中关于中程打击无人机部署的经验启示
《基于强化学习的自动化红队测试》
专知会员服务
6+阅读 · 7月23日
相关VIP内容
【ChatGPT系列报告】ChatGPT/GPT-4 如何赋能应用,31页pdf
专知会员服务
167+阅读 · 2023年4月9日
AIGC行业深度报告:ChatGPT:重新定义搜索“入口”
专知会员服务
139+阅读 · 2023年2月10日
【Facebook】人工智能基准(Benchmarking)测试再思考,55页ppt
专知会员服务
32+阅读 · 2020年12月20日
相关基金
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员