
The results are in: Which AI model is the most fallible? Persuadable? Correctible? University of Arizona researchers assessed seven different generative AI large language models, or LLMs, for these three qualities during lengthy conversations. Their work, published in Nature’s Scientific Reports, reveals intrinsic limitations that might go undetected during one-off interactions.
Among the seven LLMs tested—ChatGPT (GPT-3.5, GPT-4o and GPT-4o-mini), Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B and DeepSeek-R1—they found that:
ChatGPT 3.5 was most vulnerable to reaffirming misinformation during a conversation containing repeated false statements; Claude 3.5 Sonnet was the least.
All seven were more susceptible to misinformation on obscure top...



Recent Comments