DeepSeek tagged posts

Long AI conversations reveal misinformation vulnerabilities across seven leading chatbots

chatbots
Credit: Pavel Danilyuk from Pexels

The results are in: Which AI model is the most fallible? Persuadable? Correctible? University of Arizona researchers assessed seven different generative AI large language models, or LLMs, for these three qualities during lengthy conversations. Their work, published in Nature’s Scientific Reports, reveals intrinsic limitations that might go undetected during one-off interactions.

Among the seven LLMs tested—ChatGPT (GPT-3.5, GPT-4o and GPT-4o-mini), Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B and DeepSeek-R1—they found that:

ChatGPT 3.5 was most vulnerable to reaffirming misinformation during a conversation containing repeated false statements; Claude 3.5 Sonnet was the least.
All seven were more susceptible to misinformation on obscure top...

Read More

Number’s up: Calculators hold out against AI

In July, AI models made by Google, OpenAI and DeepSeek reached gold-level scores at the annual International Mathematical Olympiad
n July, AI models made by Google, OpenAI and DeepSeek reached gold-level scores at the annual International Mathematical Olympiad.

The humble pocket calculator may not be able to keep up with the mathematical capabilities of new technology, but it will never hallucinate.

The device’s enduring reliability equates to millions of sales each year for Japan’s Casio, which is even eyeing expansion in certain regions.

Despite lightning-speed advances in artificial intelligence, chatbots still sometimes stumble on basic addition.

In contrast, “calculators always give the correct answer,” Casio executive Tomoaki Sato told AFP.

But he conceded that calculators could one day go the way of the abacus.

“It’s undeniable that the market for personal calculators used in business is on a...

Read More