成为VIP会员查看完整内容
VIP会员码认证
首页
主题
会员
服务
注册
·
登录
稀疏自编码
关注
29
综合
百科
VIP
热门
动态
论文
精华
Exploring Sparse Autoencoders in Text-Based Causal Confounding Adjustment
Arxiv
0+阅读 · 9月1日
Global-local shrinkage priors for modeling random effects in multivariate spatial small area estimation
Arxiv
0+阅读 · 8月29日
Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
Arxiv
0+阅读 · 8月29日
REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features
Arxiv
0+阅读 · 8月28日
A Deeper Analysis of Block-Sparse Featurizers
Arxiv
0+阅读 · 8月27日
Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
Arxiv
0+阅读 · 7月27日
Endogenous Resistance to Activation Steering in Language Models
Arxiv
0+阅读 · 7月5日
LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
Arxiv
0+阅读 · 7月7日
Mapping Subnational Vulnerability to Inadequate Micronutrient Intake using a Bayesian Small Area Estimation Framework
Arxiv
0+阅读 · 8月26日
Geographically Weighted Surrogate Models for Rapid Small-Area Chronic Disease Estimation
Arxiv
0+阅读 · 7月7日
When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs
Arxiv
0+阅读 · 8月26日
Cross-Contextual Vision-Language Adaptation with LoRA for Personalized Severe Adverse Event Detection in Clinical Wound Monitoring
Arxiv
0+阅读 · 7月6日
When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control
Arxiv
0+阅读 · 7月11日
From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving
Arxiv
0+阅读 · 8月9日
The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline
Arxiv
0+阅读 · 8月24日
参考链接
提示
微信扫码
咨询专知VIP会员与技术项目合作
(加微信请备注: "专知")
微信扫码咨询专知VIP会员
Top