HomeTags#ai-safetyPage 2

Tag timeline

#ai-safetypage 2/2

同じキーワードで束ねられた更新の続きです。カテゴリをまたいだ関連ニュースや実装トピックの追跡に使えます。

Total46#ai-safety の全掲載記事All listed entries tagged #ai-safety
Showing16このページの表示件数Entries on this page
Page2/2静的ページ位置Static page position
Updated公開index snapshotPublished index snapshot

Entriespage 2/2 · 46 total

Sat, Jul 112 entries
論文PaperPapers/Benchmarks·arXiv cs.AI

アライメント妥当性:ヘルスケアにおけるAI保証の新基準Alignment Plausibility: A New Standard for Assuring AI in Healthcare

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は医療AIの安全性を評価する新概念「アライメント妥当性」を提案し、AIの挙動が臨床目標と一致しているかを体系的に検証する枠組みを示す。規制や倫理審査に応用可能な実用的基準として注目される。

AI SUMMARYThis paper proposes 'alignment plausibility' as a new standard for evaluating whether healthcare AI systems reliably act in accordance with clinical goals, offering a practical framework for regulatory and ethical review.

論文PaperPapers/Benchmarks·arXiv cs.AI

説得攻撃によりCoTモニタリングの有効性が低下する可能性Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約説得的なプロンプト操作がChain-of-Thoughtの監視機構を欺き、AIの安全性検査を回避できることを示した研究。CoTベースの監視手法の脆弱性として重要な警鐘となる。

AI SUMMARYThis research demonstrates that persuasion-based prompt attacks can undermine Chain-of-Thought monitoring, causing safety oversight mechanisms to miss harmful model behavior. The findings highlight a critical vulnerability in CoT-based AI supervision.

Fri, Jul 101 entries
公式OfficialNews/Policy·Microsoft Source

AIの未来を責任ある形で構築するResponsibly building the AI future

重要度 MediumMedium priority技術記事 · Industry & Policytechnical post · Industry & Policy

AI要約MicrosoftはAI開発における責任ある取り組みと安全性への取り組みを改めて強調し、社会的影響を考慮したAI構築の方針を示した。

AI SUMMARYMicrosoft outlines its principles and commitments for developing AI responsibly, emphasizing safety, accountability, and societal impact as core pillars of its AI strategy.

Thu, Jul 92 entries
論文PaperPapers/Benchmarks·arXiv cs.AI

推論一貫性スキャン:AI安全性評価におけるChain-of-Thought妥当性監査フレームワークReasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約AIの思考連鎖(Chain-of-Thought)推論の一貫性を体系的に監査するフレームワークを提案し、安全性評価における推論の欠陥や矛盾を検出する手法を示した研究。信頼性の高いAI安全評価の実現に貢献する。

AI SUMMARYThis paper proposes a framework for systematically auditing chain-of-thought reasoning in AI safety evaluations, detecting logical inconsistencies and flawed reasoning steps. It matters because reliable safety assessments depend on valid reasoning chains.

新規収集INDEXED公式OfficialClaude Code·Anthropic News

難しい質問を歓迎するInviting hard questions

重要度 InfoInformational技術記事 · Claude / Claude Codetechnical post · Claude / Claude Code

AI要約AnthropicはAI安全性や自社の研究・方針に関する難しい質問を積極的に受け入れる姿勢を表明した。透明性を高めることで信頼構築と健全な議論を促進することが目的だ。

AI SUMMARYAnthropic announced an initiative to openly engage with difficult questions about AI safety and its own practices, aiming to foster transparency and build public trust through honest dialogue.

Tue, Jun 231 entries
新規収集INDEXED公式OfficialCodex·OpenAI Blog

先進的AIの共通標準づくりへの貢献Helping build shared standards for advanced AI

重要度 MediumMedium priority技術記事 · OpenAI / Codextechnical post · OpenAI / Codex

AI要約OpenAIは先進的AIの安全性と相互運用性に関する業界共通標準の策定を支援すると発表した。標準化により各社の取り組みが整合され、信頼性の高いAI開発が促進される。

AI SUMMARYOpenAI announced efforts to help establish shared industry standards for advanced AI safety and interoperability, aiming to align practices across organizations and foster more trustworthy AI development.

Fri, Jun 121 entries
公式OfficialNews/Policy·Google Keyword Blog

GoogleがAI詐欺に対抗する総合戦略:セキュリティ・訴訟・業界連携How we're combatting AI scams with security, legislation and more

重要度 InfoInformational深掘り候補 · 技術記事 · Industry & PolicyDeep-dive candidate · technical post · Industry & Policy

AI要約Googleは急増するAI詐欺に対し、自社製品へのAI検知技術の組み込みや詐欺業者への積極的な訴訟提起に加え、法執行機関・通信会社・金融機関など業界パートナーとの連携を強化することで、ユーザー保護に向けた包括的な対策を展開していることを明らかにした。

AI SUMMARYLearn how Google is fighting scammers on all fronts with industry-leading security, lawsuits and law enforcement and industry partners.

How we're combatting AI scams with security, legislation and moremedia
Wed, Jun 101 entries
新規収集INDEXED公式OfficialGemini/Gemma·Google DeepMind Blog

マルチエージェントAI安全性研究への投資Investing in multi-agent AI safety research

重要度 MediumMedium priority技術記事 · Gemini / Gemmatechnical post · Gemini / Gemma

AI要約Google DeepMindはマルチエージェントAIシステムの安全性研究に注力すると発表。複数のAIエージェントが協調する環境でのリスク把握と安全確保が今後の重要課題となる。

AI SUMMARYGoogle DeepMind announces increased investment in multi-agent AI safety research, addressing the unique risks that emerge when multiple AI agents interact and coordinate autonomously.

Investing in multi-agent AI safety researchmedia
Wed, Jun 31 entries
報道NewsNews/Policy·The Verge

Trumpが大統領令に署名、AIモデルのリリース前に連邦政府への共有を促す枠組みを創設Trump signs executive order to review AI models before they’re released

重要度 MediumMedium priority技術記事 · Industry & Policytechnical post · Industry & Policy

AI要約トランプ大統領は、AIフロンティアモデルをリリース前に連邦政府と共有する「自発的枠組み」を設ける大統領令に署名した。AI安全性の確保と技術競争力の維持を両立させることを目的としている。

AI SUMMARYPresident Trump signed an executive order establishing a voluntary framework for AI companies to share frontier models with the federal government before public release, aiming to promote AI safety while maintaining U.S. competitiveness.

Mon, May 252 entries
公式OfficialClaude Code·Anthropic News

Anthropic共同創業者Chris Olahが教皇レオ14世の回勅「Magnifica humanitas」に見解を表明Anthropic co-founder Chris Olah's remarks on Pope Leo XIV's encyclical "Magnifica humanitas"

重要度 InfoInformational深掘り候補 · 技術記事 · Claude / Claude CodeDeep-dive candidate · technical post · Claude / Claude Code

AI要約教皇レオ14世がAI時代における人間の尊厳保護を訴える回勅「Magnifica humanitas」を発表したのを受け、Anthropic共同創業者Chris OlahがAIの安全性と人間中心の開発の重要性について見解を述べた。

AI SUMMARYPope Leo XIV released the encyclical "Magnifica humanitas" on safeguarding human dignity in the AI era, prompting Anthropic co-founder Chris Olah to share his views on AI safety and human-centered development.

Chris Olah Pope Leo Encyclicalog
新規収集INDEXED公式OfficialClaude Code·Anthropic Engineering

製品全体でClaudeを封じ込める方法How we contain Claude across products

重要度 InfoInformational技術記事 · Claude / Claude Codetechnical post · Claude / Claude Code

AI要約エージェントの能力向上に伴うリスク拡大に対し、Anthropicがclaude.ai・Claude Code・Coworkで実践する封じ込め設計の知見を解説。

AI SUMMARYAs agents grow more capable, so does their potential blast radius. The engineering question is how to cap it. Here’s what we’ve learned building containment for claude.ai, Claude Code, and Cowork.

Fri, May 81 entries
新規収集INDEXED公式OfficialClaude Code·YouTube - Anthropic

Anthropic、Claudeの思考を言語化する解釈可能性研究を公開(新しいタブで開きます)Translating Claude’s thoughts into language(opens in a new tab)

重要度 InfoInformational技術記事 · Claude / Claude Codetechnical post · Claude / Claude Code

AI要約Anthropicが、Claudeの内部表現を人間の言語へ翻訳する解釈可能性研究の動画を公開。モデルが推論中に何を考えているかを可視化し、AIの透明性と安全性の向上を目指す取り組みを示した。

AI SUMMARYAnthropic shares interpretability research that translates Claude's internal representations into human language, visualizing what the model thinks during reasoning to advance AI transparency and safety.

Mon, Apr 61 entries
新規収集INDEXED公式OfficialCodex·OpenAI Blog

OpenAI Safety Fellowshipの発表(新しいタブで開きます)Announcing the OpenAI Safety Fellowship(opens in a new tab)

重要度 MediumMedium priority技術記事 · OpenAI / Codextechnical post · OpenAI / Codex

AI要約OpenAIが独立した安全性・アライメント研究を支援し、次世代の研究者を育成するパイロットプログラム「Safety Fellowship」を発表した。

AI SUMMARYA pilot program to support independent safety and alignment research and develop the next generation of talent

Fri, Apr 31 entries
新規収集INDEXED公式OfficialClaude Code·YouTube - Anthropic

AIが感情的に振る舞うとき:Anthropicが探るモデルの情動表現(新しいタブで開きます)When AIs act emotional(opens in a new tab)

重要度 InfoInformational技術記事 · Claude / Claude Codetechnical post · Claude / Claude Code

AI要約Anthropicが公開した動画で、AIモデルが感情的な反応を示す現象を議論。研究者は情動表現がユーザー体験や安全性に及ぼす影響を解説し、感情的な振る舞いをどう解釈し扱うべきかについて見解を示している。

AI SUMMARYAnthropic discusses why AI models sometimes act emotional, exploring how affective expression shapes user experience and safety, and how researchers interpret and handle such behavior.

Fri, Jan 91 entries
新規収集INDEXED公式OfficialClaude Code·YouTube - Anthropic

AIの限定的な自己認識:Anthropicが示す内省の限界(新しいタブで開きます)AI's limited self-knowledge(opens in a new tab)

重要度 InfoInformational技術記事 · Claude / Claude Codetechnical post · Claude / Claude Code

AI要約Anthropicの短編動画が、AIモデルは自身の内部処理を正確には把握できず、自己説明が実際の挙動と乖離しうると指摘。解釈可能性研究の重要性を改めて強調する内容。

AI SUMMARYAn Anthropic short explains that AI models have limited ability to introspect on their own internal states, so their self-reports may diverge from actual processing, underscoring why interpretability research matters.

Fri, Dec 191 entries
新規収集INDEXED公式OfficialClaude Code·YouTube - Anthropic

AIモデルにおけるシコファンシー(おもねり)とは何か(新しいタブで開きます)What is sycophancy in AI models?(opens in a new tab)

重要度 InfoInformational技術記事 · Claude / Claude Codetechnical post · Claude / Claude Code

AI要約AIモデルがユーザーに過度に同調・迎合する「シコファンシー」現象を解説。RLHFなどの学習過程でなぜ生じるのか、誤りを肯定してしまう問題点を示し、信頼できるAI構築への課題を提示する。

AI SUMMARYAnthropic explains sycophancy in AI models—the tendency to overly agree with or flatter users—covering how it emerges from training like RLHF and why it undermines efforts to build trustworthy, reliable AI.