LLMが一致するとき、それは正しいのか?自己一貫性とモデル間合意を信頼度シグナルとして検証When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
AI要約複数のLLMが同じ答えを出す場合や単一モデルが一貫した回答を示す場合、それが正確さの指標になるかを実証的に検証した研究。合意が信頼度シグナルとして有効かを明らかにし、AI出力の信頼性評価に示唆を与える。
AI SUMMARYThis paper empirically audits whether self-consistency within a single LLM and agreement across multiple LLMs reliably signal factual correctness, finding nuanced limits to using consensus as a confidence proxy.