Here’s why ChaptGPT and other AI models may not always improve over time

Here’s why ChaptGPT and other AI models may not always improve over time
redit: Unsplash/ Andrew Neel

When OpenAI released its latest text-generating artificial intelligence, the large language model GPT-4, in March, it was very good at identifying prime numbers. When the AI was given a series of 500 such numbers and asked whether they were primes, it correctly labeled them 97.6 percent of the time. But a few months later, in June, the same test yielded very different results. GPT-4 only correctly labeled 2.4 percent of the prime numbers AI researchers prompted it with—a complete reversal in apparent accuracy. The finding underscores the complexity of large artificial intelligence models: instead of AI uniformly improving at every task on a straight trajectory, the reality is much more like a winding road full of speed bumps and detours.

Mis- and disinformation are distorting science and public policy. Subscribe to our Daily and Weekly Digests for incisive coverage of ‘disruptive’ innovations in AI, agricultural biotechnology, food, chemicals, nuclear energy, vaccines, and other disruptive innovations.

Even OpenAI has acknowledged that, when it comes to GPT-4, “while the majority of metrics have improved, there may be some tasks where the performance gets worse,” as employees of the company wrote in a July 20 update to a post on OpenAi’s blog. Past studies of other models have also shown this sort of behavioral shift, or “model drift,” over time. That alone could be a big problem for developers and researchers who’ve come to rely on this AI in their own work.

This is an excerpt. Read the full article here

Related Articles

Infographic: Global regulatory and health research agencies on whether glyphosate causes cancer

Infographic: Global regulatory and health research agencies on whether glyphosate causes cancer

Does glyphosate—the world's most heavily-used herbicide—pose serious harm to humans? Is it carcinogenic? Those issues are of both legal and ...

Most Popular

ChatGPT-Image-Mar-10-2026-01_39_01-PM
Viewpoint—“Miracle molecule” debunked: Why acemannan supplements don’t work
Screenshot-2026-10-01-at-10.52.08-AM
Non-browning CRISPR gene-edited bananas move one step closer to British markets
AI-Force-Leadership-Portrait
Jay Clayton controversy: The divisive history of Trump’s newly-appointed AI czar
Screenshot-2026-09-22-at-10.54.12-AM
‘They misunderstand science’: Research attempting to tie autism to changes in the gut microbiome comes under fire
Elderly-Speaker-at-MAHA-Institute
At anti-vaccine MAHA conference, RFK Jr. pledges to unify federal health information and combat vaccine safety ‘misinformation’
Screenshot-2026-09-28-at-11.26.17-AM
Viewpoint: Wellness movement activists reposition nicotine from a killer addictive drug into a cure-all for brain ills
ChatGPT-Image-Sep-21-2026-11_21_39-AM
Viewpoint: Food cults: Online ‘nutrition communities’ can pose unique dangers
Screenshot 2026-10-01 at 12.30
Combating ‘narrative distortion’: Will pharma companies step up to challenge health misinformation?
Screenshot-2026-10-01-at-11.21.27-AM
Vaccine hesitancy, like malaria, surges globally, topping 35% in some countries
ChatGPT-Image-Sep-21-2026-03_15_09-PM
Viewpoint: Chipotle-inspired lesson: 10,000 sick customers suggest that processed foods can be safer than fresh ones
generated-1790707585414
Facts & Fallacies podcast: America's health care system is broken. A doctor explains how to fix it
ChatGPT-Image-Sep-30-2026-01_38_36-PM
Viewpoint: Died “with measles” vs. died “from measles”: RFK, Jr.’s dangerously subversive word games
glp menu logo outlined

Get news on human & agricultural genetics and biotechnology delivered to your inbox.