* . *
  • About
  • Advertise
  • Privacy & Policy
  • Contact
Friday, August 15, 2025
Earth-News
  • Home
  • Business
  • Entertainment
    I’ll miss the chaos of ‘And Just like That…’ (and Che Diaz too) – yahoo.com

    Why I’ll Truly Miss the Wild Ride of ‘And Just Like That…’ (and Che Diaz!)

    Webtoon Entertainment Stages Recovery With Disney’s Stamp of Approval – The Wall Street Journal

    Webtoon Entertainment Soars to New Heights with Disney’s Stamp of Approval

    Georgia Tech Launches Arts, Entertainment, and Creative Technologies Degree – Georgia Tech News Center

    Georgia Tech Unveils Exciting New Degree in Arts, Entertainment, and Creative Technologies

    John Davison departs from IGN Entertainment – GamesIndustry.biz

    John Davison Steps Down from IGN Entertainment Leadership

    JPMorgan raises Flutter Entertainment stock price target to GBP273 – Investing.com

    JPMorgan Raises Flutter Entertainment Price Target to £273, Signaling Strong Growth Ahead

    Star Entertainment reaches deal to sell 50% stake in Brisbane resort to HK investors – Reuters

    Star Entertainment Seals Landmark Deal, Sells Half of Brisbane Resort to Hong Kong Investors

  • General
  • Health
  • News

    Cracking the Code: Why China’s Economic Challenges Aren’t Shaking Markets, Unlike America’s” – Bloomberg

    Trump’s Narrow Window to Spread the Truth About Harris

    Trump’s Narrow Window to Spread the Truth About Harris

    Israel-Gaza war live updates: Hamas leader Ismail Haniyeh assassinated in Iran, group says

    Israel-Gaza war live updates: Hamas leader Ismail Haniyeh assassinated in Iran, group says

    PAP Boss to Niger Delta Youths, Stay Away from the Protest

    PAP Boss to Niger Delta Youths, Stay Away from the Protest

    Court Restricts Protests In Lagos To Freedom, Peace Park

    Court Restricts Protests In Lagos To Freedom, Peace Park

    Fans React to Jazz Jennings’ Inspiring Weight Loss Journey

    Fans React to Jazz Jennings’ Inspiring Weight Loss Journey

    Trending Tags

    • Trump Inauguration
    • United Stated
    • White House
    • Market Stories
    • Election Results
  • Science
  • Sports
  • Technology
    Verb Technology Reports Revenue Growth Amidst Strategic Expansions – TipRanks

    Verb Technology Soars with Impressive Revenue Growth Driven by Strategic Expansions

    Midwest Technology Summit held in Fargo – WDAY Radio

    Midwest Technology Summit held in Fargo – WDAY Radio

    K1 Semiconductor Joins Chicago Quantum Exchange To Advance Wafer Technology. – Quantum Zeitgeist

    K1 Semiconductor Partners with Chicago Quantum Exchange to Revolutionize Wafer Technology

    Indirect tax transformation: Navigating change, embracing technology – Thomson Reuters tax and accounting

    Revolutionizing Indirect Tax: Embracing Technology to Navigate Change

    California’s wildfire moonshot: How new technology will defeat advancing flames – Los Angeles Times

    California’s Wildfire Revolution: How Cutting-Edge Technology Is Poised to Stop Raging Flames

    LSU grad uses 3D printing to create adaptive technology for children – CBS News

    LSU Graduate Revolutionizes Adaptive Technology for Kids with 3D Printing

    Trending Tags

    • Nintendo Switch
    • CES 2017
    • Playstation 4 Pro
    • Mark Zuckerberg
No Result
View All Result
  • Home
  • Business
  • Entertainment
    I’ll miss the chaos of ‘And Just like That…’ (and Che Diaz too) – yahoo.com

    Why I’ll Truly Miss the Wild Ride of ‘And Just Like That…’ (and Che Diaz!)

    Webtoon Entertainment Stages Recovery With Disney’s Stamp of Approval – The Wall Street Journal

    Webtoon Entertainment Soars to New Heights with Disney’s Stamp of Approval

    Georgia Tech Launches Arts, Entertainment, and Creative Technologies Degree – Georgia Tech News Center

    Georgia Tech Unveils Exciting New Degree in Arts, Entertainment, and Creative Technologies

    John Davison departs from IGN Entertainment – GamesIndustry.biz

    John Davison Steps Down from IGN Entertainment Leadership

    JPMorgan raises Flutter Entertainment stock price target to GBP273 – Investing.com

    JPMorgan Raises Flutter Entertainment Price Target to £273, Signaling Strong Growth Ahead

    Star Entertainment reaches deal to sell 50% stake in Brisbane resort to HK investors – Reuters

    Star Entertainment Seals Landmark Deal, Sells Half of Brisbane Resort to Hong Kong Investors

  • General
  • Health
  • News

    Cracking the Code: Why China’s Economic Challenges Aren’t Shaking Markets, Unlike America’s” – Bloomberg

    Trump’s Narrow Window to Spread the Truth About Harris

    Trump’s Narrow Window to Spread the Truth About Harris

    Israel-Gaza war live updates: Hamas leader Ismail Haniyeh assassinated in Iran, group says

    Israel-Gaza war live updates: Hamas leader Ismail Haniyeh assassinated in Iran, group says

    PAP Boss to Niger Delta Youths, Stay Away from the Protest

    PAP Boss to Niger Delta Youths, Stay Away from the Protest

    Court Restricts Protests In Lagos To Freedom, Peace Park

    Court Restricts Protests In Lagos To Freedom, Peace Park

    Fans React to Jazz Jennings’ Inspiring Weight Loss Journey

    Fans React to Jazz Jennings’ Inspiring Weight Loss Journey

    Trending Tags

    • Trump Inauguration
    • United Stated
    • White House
    • Market Stories
    • Election Results
  • Science
  • Sports
  • Technology
    Verb Technology Reports Revenue Growth Amidst Strategic Expansions – TipRanks

    Verb Technology Soars with Impressive Revenue Growth Driven by Strategic Expansions

    Midwest Technology Summit held in Fargo – WDAY Radio

    Midwest Technology Summit held in Fargo – WDAY Radio

    K1 Semiconductor Joins Chicago Quantum Exchange To Advance Wafer Technology. – Quantum Zeitgeist

    K1 Semiconductor Partners with Chicago Quantum Exchange to Revolutionize Wafer Technology

    Indirect tax transformation: Navigating change, embracing technology – Thomson Reuters tax and accounting

    Revolutionizing Indirect Tax: Embracing Technology to Navigate Change

    California’s wildfire moonshot: How new technology will defeat advancing flames – Los Angeles Times

    California’s Wildfire Revolution: How Cutting-Edge Technology Is Poised to Stop Raging Flames

    LSU grad uses 3D printing to create adaptive technology for children – CBS News

    LSU Graduate Revolutionizes Adaptive Technology for Kids with 3D Printing

    Trending Tags

    • Nintendo Switch
    • CES 2017
    • Playstation 4 Pro
    • Mark Zuckerberg
No Result
View All Result
Earth-News
No Result
View All Result
Home Technology

Boffins find AI stumbles when quizzed on the tough stuff

October 30, 2023
in Technology
Boffins find AI stumbles when quizzed on the tough stuff
Share on FacebookShare on Twitter

AI models can manage well enough when prompted with text or images, and may even solve complex problems when not making terrible errors.

OpenAI, for example, has said that its GPT-4 model managed to score 700 out of 800 on the SAT math exam. Not all such claims have borne out, however: A paper released in June that said GPT-4 could get a computer science degree at MIT was subsequently withdrawn.

So to better assess how large language models – which interpret text input – and large multimodal models – which interpret text, images and perhaps other forms of input – actually handle problem solving, a group of ten researchers from the University of California, Los Angeles, the University of Washington, and Microsoft Research have devised a testing benchmark called MathVista that focuses on visually-oriented challenges.

“The ability of these foundation models to perform mathematical reasoning in visual contexts has not been systematically examined,” say the authors – Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao, in a preprint paper [PDF].

It is thus essential, they say, to develop a new benchmark to help the development of mathematical reasoning with a visual component and to evaluate how various models compare at reasoning tasks.

Being able to show that one’s AI model can correctly solve visual problems may prove helpful in determining whether it’s wise to, say, trust software to drive a car without stopping atop an accident victim.

MathVista incorporates 6,141 examples that were developed from 28 multimodal datasets and from 3 new datasets called IQTest, FunctionQA, and PaperQA. It covers various forms of reasoning (algebraic, arithmetic, geometric, logical, numeric, scientific, and statistical), with a focus on figure question answering, geometry problem solving, math word problems, textbook questions, and visual questions.

Screenshot of MathVista challenge question

Screenshot of MathVista challenge question – Click to enlarge

The researchers tested a dozen foundation models: three LLMs ChatGPT, GPT-4, and Claude-2), two proprietary LMMs (GPT4V and Bard), and seven open-source LMMs. They also considered human answers, provided via Amazon Mechanical Turkers with at least a high school degree, and random responses.

AWS CEO talks up AI to focus minds of Wall Street types

Clippy-like AI at forefront of Windows update previews

Bug bounty hunters load up to stalk AI and fancy bagging big bucks

How prompt injection attacks hijack today’s top-end AI – and it’s tough to fix

The good news for AI practitioners is that the LLMs and LMMs all did better than random chance, which isn’t all that surprising considering that many of the questions were multiple choice rather than yes or no.

In fact, the top performer, OpenAI’s GPT-4V, managed to surpass human performance in specific areas – questions involving algebraic reasoning and complex visual challenges involving tables and function plots.

We note that Microsoft, whose researchers contributed to this project, has a substantial stake in OpenAI.

The less good news is that even GPT-4V only managed to get 49.9 percent of the questions correct. That’s adequate if the goal is to best multimodal Bard, which managed an accuracy percentage of 34.8 percent.

But it’s still shy of the Amazon Mechanical Turk workers who were put to the test and managed a score of 60.3 percent. As the researchers observe in their paper, “a 10.4 percent gap in overall accuracy remains when compared to the human baseline, leaving plenty of room for model improvement.” ®

>>> Read full article>>>
Copyright for syndicated content belongs to the linked Source : The Register – https://go.theregister.com/feed/www.theregister.com/2023/10/29/ai_math_quiz/

Tags: Boffinsstumblestechnology
Previous Post

Tenfold electric vehicles on 2030 roads could be a shock to the system

Next Post

Somehow, The Black Phone Will Ring Up a Sequel

Chinese premier urges efforts to write new chapter in building ecological civilization in new era – China Daily

Chinese premier urges efforts to write new chapter in building ecological civilization in new era – China Daily

August 15, 2025
‘Like a creeping mold that’s spreading across the landscape’: Separate dry areas around the world are merging into ‘mega-drying’ regions at an alarming rate, study finds – Live Science

‘Like a creeping mold that’s spreading across the landscape’: Separate dry areas around the world are merging into ‘mega-drying’ regions at an alarming rate, study finds – Live Science

August 15, 2025
A new color of grief – Lifestyle.INQ

Unveiling a New Shade of Grief: Embracing Unexpected Emotions

August 15, 2025
Verb Technology Reports Revenue Growth Amidst Strategic Expansions – TipRanks

Verb Technology Soars with Impressive Revenue Growth Driven by Strategic Expansions

August 15, 2025
Lakers honoring legendary coach Pat Riley with statue to be unveiled on Feb. 22 – Yahoo Sports

Lakers to Reveal Breathtaking Statue Honoring Legendary Coach Pat Riley on February 22

August 15, 2025
Is This the Hardest Physical Contest in the World? – The Atlantic

Could This Be the Toughest Physical Challenge on Earth?

August 15, 2025
Trump’s tariffs, other federal policies starting to ding Minnesota’s economy – Star Tribune

Trump’s tariffs, other federal policies starting to ding Minnesota’s economy – Star Tribune

August 15, 2025
I’ll miss the chaos of ‘And Just like That…’ (and Che Diaz too) – yahoo.com

Why I’ll Truly Miss the Wild Ride of ‘And Just Like That…’ (and Che Diaz!)

August 15, 2025
How parents can help support college students’ mental health – WALB

How Parents Can Play a Vital Role in Supporting Their College Students’ Mental Health

August 15, 2025
‘Tesla shame’ bypasses Norway as sales jump despite Musk’s politics – Yahoo Finance

Tesla Sales Skyrocket in Norway Amidst Controversy Over Musk’s Politics

August 15, 2025

Categories

Archives

August 2025
MTWTFSS
 123
45678910
11121314151617
18192021222324
25262728293031
« Jul    
Earth-News.info

The Earth News is an independent English-language daily published Website from all around the World News

Browse by Category

  • Business (20,132)
  • Ecology (772)
  • Economy (794)
  • Entertainment (21,671)
  • General (16,478)
  • Health (9,833)
  • Lifestyle (805)
  • News (22,149)
  • People (796)
  • Politics (802)
  • Science (16,007)
  • Sports (21,292)
  • Technology (15,774)
  • World (777)

Recent News

Chinese premier urges efforts to write new chapter in building ecological civilization in new era – China Daily

Chinese premier urges efforts to write new chapter in building ecological civilization in new era – China Daily

August 15, 2025
‘Like a creeping mold that’s spreading across the landscape’: Separate dry areas around the world are merging into ‘mega-drying’ regions at an alarming rate, study finds – Live Science

‘Like a creeping mold that’s spreading across the landscape’: Separate dry areas around the world are merging into ‘mega-drying’ regions at an alarming rate, study finds – Live Science

August 15, 2025
  • About
  • Advertise
  • Privacy & Policy
  • Contact

© 2023 earth-news.info

No Result
View All Result

© 2023 earth-news.info

No Result
View All Result

© 2023 earth-news.info

Go to mobile version