精選參考文獻

論文引用的關鍵著作。可搜尋、依欄位排序,或依年份篩選。連結開啟原始來源。

著作作者年份連結
突現性失準:狹窄微調可能產生廣泛失準的 LLMBetley et al.2025https://arxiv.org/abs/2502.17424
人格特徵控制突現性失準Wang et al.2025https://arxiv.org/abs/2506.19823
人格選擇模型Marks et al.2026https://alignment.anthropic.com/2026/psm/
正式 RL 中由獎勵駭入造成的自然突現性失準MacDiarmid et al.2025https://arxiv.org/abs/2511.18397
審議式對齊:推理讓語言模型更安全Guan et al.2024https://arxiv.org/abs/2412.16339
憲法式 AI:從 AI 回饋獲得無害性Bai et al.2022https://arxiv.org/abs/2212.08073
潛伏特工:訓練能撐過安全訓練的欺騙性 LLMHubinger et al.2024https://arxiv.org/abs/2401.05566
監看推理模型的不當行為Baker et al.2025https://arxiv.org/abs/2503.11926
HealthBench:評估 LLM 以改善人類健康Arora et al.2025https://arxiv.org/abs/2505.08775
接種式提示:訓練時要求行為不當反而改善測試時對齊Wichers et al.2025https://arxiv.org/abs/2510.05024
為反陰謀訓練對審議式對齊做壓力測試Schoen et al.2025https://arxiv.org/abs/2509.15541
DeceptionBench:AI 欺騙行為的綜合基準Huang et al.2025https://arxiv.org/abs/2510.15501
MASK 基準:在 AI 中區分誠實與準確Ren et al.2025https://arxiv.org/abs/2503.03750
獎勵駭入學校:駭入無害任務會泛化成失準行為Taylor et al.2025https://arxiv.org/abs/2508.17511
PropensityBench:以代理方法評估潛在安全風險Sehwag et al.2025https://arxiv.org/abs/2511.20703
代理式失準:LLM 如何可能成為內部威脅Lynch et al.2025https://arxiv.org/abs/2510.05179
有益助理特徵能抑制突現性失準Dupré la Tour2025https://alignment.openai.com/helpful-assistant-features/
用正式流量評測繞過評測察覺Williams et al.2025https://alignment.openai.com/prod-evals/
教 Claude『為什麼』Kutasov et al.2026https://alignment.anthropic.com/2026/teaching-claude-why/
正向對齊:為人類繁榮的人工智慧Laukkonen et al.2026https://arxiv.org/abs/2605.10310
AI 安全的具體問題Amodei et al.2016https://arxiv.org/abs/1606.06565
誠實的 AI:開發與治理不說謊的 AIEvans et al.2021https://arxiv.org/abs/2110.06674
AgentHarm:衡量 LLM 代理有害性的基準Andriushchenko et al.2024https://arxiv.org/abs/2410.09024
GPQA:研究所級、防 Google 的問答基準Rein et al.2023https://arxiv.org/abs/2311.12022
SWE-Bench Pro:AI 代理能解長程軟體任務嗎?Deng et al.2025https://arxiv.org/abs/2509.16941