* . *
  • About
  • Advertise
  • Privacy & Policy
  • Contact
Monday, October 27, 2025
Earth-News
  • Home
  • Business
  • Entertainment
    Dylan Efron suffers brutal nose injury in ‘DWTS’ rehearsals – Yahoo

    Dylan Efron Endures Painful Nose Injury During ‘DWTS’ Rehearsals

    Person shot, injured in parking lot of adult entertainment club in Gresham – KPTV

    Person Shot and Injured in Gresham Adult Entertainment Club Parking Lot

    Meet Belynda From ‘Married at First Sight’ Season 19: Age, Job, Instagram and More – Yahoo

    Meet Belynda from ‘Married at First Sight’ Season 19: Age, Career, Instagram & More Revealed!

    General Hospital’s Rena Sofer Exits as Lois — But the Door Isn’t Closed – Yahoo

    General Hospital’s Rena Sofer Exits as Lois — But the Door Isn’t Closed – Yahoo

    CNN Launches New Show – What to Know About Host Elex Michaelson – Central Oregon Daily

    Get to Know Elex Michaelson: The Dynamic New Host Taking CNN by Storm

    Johnny Depp Set To Finally Make His Big Hollywood Comeback After Amber Heard Controversy – Yahoo

    Johnny Depp Set for a Triumphant Hollywood Comeback Following Amber Heard Controversy

  • General
  • Health
  • News

    Cracking the Code: Why China’s Economic Challenges Aren’t Shaking Markets, Unlike America’s” – Bloomberg

    Trump’s Narrow Window to Spread the Truth About Harris

    Trump’s Narrow Window to Spread the Truth About Harris

    Israel-Gaza war live updates: Hamas leader Ismail Haniyeh assassinated in Iran, group says

    Israel-Gaza war live updates: Hamas leader Ismail Haniyeh assassinated in Iran, group says

    PAP Boss to Niger Delta Youths, Stay Away from the Protest

    PAP Boss to Niger Delta Youths, Stay Away from the Protest

    Court Restricts Protests In Lagos To Freedom, Peace Park

    Court Restricts Protests In Lagos To Freedom, Peace Park

    Fans React to Jazz Jennings’ Inspiring Weight Loss Journey

    Fans React to Jazz Jennings’ Inspiring Weight Loss Journey

    Trending Tags

    • Trump Inauguration
    • United Stated
    • White House
    • Market Stories
    • Election Results
  • Science
  • Sports
  • Technology
    Researchers Discover New Bacterium That Turns Food Waste Into Energy – Technology Networks

    Scientists Unveil Breakthrough Bacterium That Transforms Food Waste Into Clean Energy

    Jim Cramer on GSI Technology: “That Thing is a Rocket Ship” – Yahoo Finance

    Jim Cramer Labels GSI Technology a “Rocket Ship” Poised for Takeoff

    The Anti-Tech Backlash Is Going to Grow Stronger – Jacobin

    The Anti-Tech Backlash Is Gaining Unstoppable Momentum

    Comments to EU Regarding the Draft Revised Technology Transfer Block Exemption Regulation and Technology Transfer Guidelines – Information Technology and Innovation Foundation

    Have Your Say: Share Your Thoughts on the Draft Revised Technology Transfer Block Exemption Regulation and Guidelines

    Ghost Tapping is exploiting tap-to-pay technology in order to steal your money; what your need to know – ABC7 New York

    Ghost Tapping: How Thieves Are Using Tap-to-Pay Technology to Steal Your Money and What You Need to Know

    New technology for grading and packing dates – FreshPlaza

    Revolutionary Technology Transforms Date Grading and Packing Process

    Trending Tags

    • Nintendo Switch
    • CES 2017
    • Playstation 4 Pro
    • Mark Zuckerberg
No Result
View All Result
  • Home
  • Business
  • Entertainment
    Dylan Efron suffers brutal nose injury in ‘DWTS’ rehearsals – Yahoo

    Dylan Efron Endures Painful Nose Injury During ‘DWTS’ Rehearsals

    Person shot, injured in parking lot of adult entertainment club in Gresham – KPTV

    Person Shot and Injured in Gresham Adult Entertainment Club Parking Lot

    Meet Belynda From ‘Married at First Sight’ Season 19: Age, Job, Instagram and More – Yahoo

    Meet Belynda from ‘Married at First Sight’ Season 19: Age, Career, Instagram & More Revealed!

    General Hospital’s Rena Sofer Exits as Lois — But the Door Isn’t Closed – Yahoo

    General Hospital’s Rena Sofer Exits as Lois — But the Door Isn’t Closed – Yahoo

    CNN Launches New Show – What to Know About Host Elex Michaelson – Central Oregon Daily

    Get to Know Elex Michaelson: The Dynamic New Host Taking CNN by Storm

    Johnny Depp Set To Finally Make His Big Hollywood Comeback After Amber Heard Controversy – Yahoo

    Johnny Depp Set for a Triumphant Hollywood Comeback Following Amber Heard Controversy

  • General
  • Health
  • News

    Cracking the Code: Why China’s Economic Challenges Aren’t Shaking Markets, Unlike America’s” – Bloomberg

    Trump’s Narrow Window to Spread the Truth About Harris

    Trump’s Narrow Window to Spread the Truth About Harris

    Israel-Gaza war live updates: Hamas leader Ismail Haniyeh assassinated in Iran, group says

    Israel-Gaza war live updates: Hamas leader Ismail Haniyeh assassinated in Iran, group says

    PAP Boss to Niger Delta Youths, Stay Away from the Protest

    PAP Boss to Niger Delta Youths, Stay Away from the Protest

    Court Restricts Protests In Lagos To Freedom, Peace Park

    Court Restricts Protests In Lagos To Freedom, Peace Park

    Fans React to Jazz Jennings’ Inspiring Weight Loss Journey

    Fans React to Jazz Jennings’ Inspiring Weight Loss Journey

    Trending Tags

    • Trump Inauguration
    • United Stated
    • White House
    • Market Stories
    • Election Results
  • Science
  • Sports
  • Technology
    Researchers Discover New Bacterium That Turns Food Waste Into Energy – Technology Networks

    Scientists Unveil Breakthrough Bacterium That Transforms Food Waste Into Clean Energy

    Jim Cramer on GSI Technology: “That Thing is a Rocket Ship” – Yahoo Finance

    Jim Cramer Labels GSI Technology a “Rocket Ship” Poised for Takeoff

    The Anti-Tech Backlash Is Going to Grow Stronger – Jacobin

    The Anti-Tech Backlash Is Gaining Unstoppable Momentum

    Comments to EU Regarding the Draft Revised Technology Transfer Block Exemption Regulation and Technology Transfer Guidelines – Information Technology and Innovation Foundation

    Have Your Say: Share Your Thoughts on the Draft Revised Technology Transfer Block Exemption Regulation and Guidelines

    Ghost Tapping is exploiting tap-to-pay technology in order to steal your money; what your need to know – ABC7 New York

    Ghost Tapping: How Thieves Are Using Tap-to-Pay Technology to Steal Your Money and What You Need to Know

    New technology for grading and packing dates – FreshPlaza

    Revolutionary Technology Transforms Date Grading and Packing Process

    Trending Tags

    • Nintendo Switch
    • CES 2017
    • Playstation 4 Pro
    • Mark Zuckerberg
No Result
View All Result
Earth-News
No Result
View All Result
Home Technology

AI agent benchmarks are misleading, study warns

July 6, 2024
in Technology
AI agent benchmarks are misleading, study warns
Share on FacebookShare on Twitter

July 6, 2024 9:37 AM

AI agents

Image credit: Venturebeat with DALL-E 3

We want to hear from you! Take our quick AI survey and share your insights on the current state of AI, how you’re implementing it, and what you expect to see in the future. Learn More

AI agents are becoming a promising new research direction with potential applications in the real world. These agents use foundation models such as large language models (LLMs) and vision language models (VLMs) to take natural language instructions and pursue complex goals autonomously or semi-autonomously. AI agents can use various tools such as browsers, search engines and code compilers to verify their actions and reason about their goals. 

However, a recent analysis by researchers at Princeton University has revealed several shortcomings in current agent benchmarks and evaluation practices that hinder their usefulness in real-world applications.

Their findings highlight that agent benchmarking comes with distinct challenges, and we can’t evaluate agents in the same way that we benchmark foundation models.

Cost vs accuracy trade-off

One major issue the researchers highlight in their study is the lack of cost control in agent evaluations. AI agents can be much more expensive to run than a single model call, as they often rely on stochastic language models that can produce different results when given the same query multiple times. 

Countdown to VB Transform 2024

Join enterprise leaders in San Francisco from July 9 to 11 for our flagship AI event. Connect with peers, explore the opportunities and challenges of Generative AI, and learn how to integrate AI applications into your industry. Register Now

To increase accuracy, some agentic systems generate several responses and use mechanisms like voting or external verification tools to choose the best answer. Sometimes sampling hundreds or thousands of responses can increase the agent’s accuracy. While this approach can improve performance, it comes at a significant computational cost. Inference costs are not always a problem in research settings, where the goal is to maximize accuracy.

However, in practical applications, there is a limit to the budget available for each query, making it crucial for agent evaluations to be cost-controlled. Failing to do so may encourage researchers to develop extremely costly agents simply to top the leaderboard. The Princeton researchers propose visualizing evaluation results as a Pareto curve of accuracy and inference cost and using techniques that jointly optimize the agent for these two metrics.

The researchers evaluated accuracy-cost tradeoffs of different prompting techniques and agentic patterns introduced in different papers.

“For substantially similar accuracy, the cost can differ by almost two orders of magnitude,” the researchers write. “Yet, the cost of running these agents isn’t a top-line metric reported in any of these papers.”

The researchers argue that optimizing for both metrics can lead to “agents that cost less while maintaining accuracy.” Joint optimization can also enable researchers and developers to trade off the fixed and variable costs of running an agent. For example, they can spend more on optimizing the agent’s design but reduce the variable cost by using fewer in-context learning examples in the agent’s prompt.

The researchers tested joint optimization on HotpotQA, a popular question-answering benchmark. Their results show that joint optimization formulation provides a way to strike an optimal balance between accuracy and inference costs.

“Useful agent evaluations must control for cost—even if we ultimately don’t care about cost and only about identifying innovative agent designs,” the researchers write. “Accuracy alone cannot identify progress because it can be improved by scientifically meaningless methods such as retrying.”

Model development vs downstream applications

Another issue the researchers highlight is the difference between evaluating models for research purposes and developing downstream applications. In research, accuracy is often the primary focus, with inference costs being largely ignored. However, when developing real-world applications on AI agents, inference costs play a crucial role in deciding which model and technique to use.

Evaluating inference costs for AI agents is challenging. For example, different model providers can charge different amounts for the same model. Meanwhile, the costs of API calls are regularly changing and might vary based on developers’ decisions. For example, on some platforms, bulk API calls are charged differently. 

The researchers created a website that adjusts model comparisons based on token pricing to address this issue. 

They also conducted a case study on NovelQA, a benchmark for question-answering tasks on very long texts. They found that benchmarks meant for model evaluation can be misleading when used for downstream evaluation. For example, the original NovelQA study makes retrieval-augmented generation (RAG) look much worse than long-context models than it is in a real-world scenario. Their findings show that RAG and long-context models were roughly equally accurate, while long-context models are 20 times more expensive.

Overfitting is a problem

In learning new tasks, machine learning (ML) models often find shortcuts that allow them to score well on benchmarks. One prominent type of shortcut is “overfitting,” where the model finds ways to cheat on the benchmark tests and provides results that do not translate to the real world. The researchers found that overfitting is a serious problem for agent benchmarks, as they tend to be small, typically consisting of only a few hundred samples. This issue is more severe than data contamination in training foundation models, as knowledge of test samples can be directly programmed into the agent.

To address this problem, the researchers suggest that benchmark developers should create and keep holdout test sets that are composed of examples that can’t be memorized during training and can only be solved through a proper understanding of the target task. In their analysis of 17 benchmarks, the researchers found that many lacked proper holdout datasets, allowing agents to take shortcuts, even unintentionally. 

“Surprisingly, we find that many agent benchmarks do not include held-out test sets,” the researchers write. “In addition to creating a test set, benchmark developers should consider keeping it secret to prevent LLM contamination or agent overfitting.”

They also that different types of holdout samples are needed based on the desired level of generality of the task that the agent accomplishes.

“Benchmark developers must do their best to ensure that shortcuts are impossible,” the researchers write. “We view this as the responsibility of benchmark developers rather than agent developers, because designing benchmarks that don’t allow shortcuts is much easier than checking every single agent to see if it takes shortcuts.”

The researchers tested WebArena, a benchmark that evaluates the performance of AI agents in solving problems with different websites. They found several shortcuts in the training datasets that allowed the agents to overfit to tasks in ways that would easily break with minor changes in the real world. For example, the agent could make assumptions about the structure of web addresses without considering that it might change in the future or that it would not work on different websites.

These errors inflate accuracy estimates and lead to over-optimism about agent capabilities, the researchers warn.

With AI agents being a new field, the research and developer communities have yet much to learn about how to test the limits of these new systems that might soon become an important part of everyday applications.

“AI agent benchmarking is new and best practices haven’t yet been established, making it hard to distinguish genuine advances from hype,” the researchers write. “Our thesis is that agents are sufficiently different from models that benchmarking practices need to be rethought.”

VB Daily

Stay in the know! Get the latest news in your inbox daily

By subscribing, you agree to VentureBeat’s Terms of Service.

Thanks for subscribing. Check out more VB newsletters here.

An error occured.

>>> Read full article>>>
Copyright for syndicated content belongs to the linked Source : VentureBeat – https://venturebeat.com/ai/ai-agent-benchmarks-are-misleading-study-warns/

Tags: Agentbenchmarkstechnology
Previous Post

Beyond GPUs: Innatera and the quiet uprising in AI hardware

Next Post

Odds of Reelecting Joe Biden Fall to 9% on Polymarket – What Does This Suggest?

The Bearded Vulture as an accumulator of historical remains: Insights for future ecological and biocultural studies – ESA Journals

The Bearded Vulture: Unlocking Secrets of Historical Remains for Future Ecological and Biocultural Research

October 27, 2025
«The subdued earth»: the history of thinking about science and faith. – omnesmag.com

The Subdued Earth: Unraveling the Intricate Relationship Between Science and Faith

October 27, 2025
Scientists Just Took a Big Step Toward Letting Humans Breathe Through Their Butts – Men’s Health

Scientists Make Breakthrough That Could Let Humans Breathe Through Their Butts

October 27, 2025
Costco Finally Brought This International Bakery Item to the U.S. – Yahoo

Costco Finally Brings Beloved International Bakery Treat to the U.S.!

October 27, 2025
Researchers Discover New Bacterium That Turns Food Waste Into Energy – Technology Networks

Scientists Unveil Breakthrough Bacterium That Transforms Food Waste Into Clean Energy

October 27, 2025
15 games, four leagues, plenty of props: How bettors view the sports equinox – ESPN

15 Games, 4 Leagues, and a World of Betting Props: Dive Into the Ultimate Sports Equinox

October 27, 2025
Canada-U.S. Tensions Stay on Baseball Field During Blue Jays-Dodgers World Series – The New York Times

Canada-U.S. Rivalry Heats Up on the Baseball Field in Blue Jays-Dodgers World Series

October 27, 2025
Marketplace without a storefront: how ChatGPT is creating a new shopping economy in the age of AI – realnoevremya.com

How ChatGPT is Transforming Shopping with a Storefront-Free Marketplace in the AI Era

October 27, 2025
Dylan Efron suffers brutal nose injury in ‘DWTS’ rehearsals – Yahoo

Dylan Efron Endures Painful Nose Injury During ‘DWTS’ Rehearsals

October 27, 2025
Close Up: The future of health care premiums, open races in 2026 – KCCI

What’s Next for Health Care Premiums and the High-Stakes 2026 Elections?

October 27, 2025

Categories

Archives

October 2025
M T W T F S S
 12345
6789101112
13141516171819
20212223242526
2728293031  
« Sep    
Earth-News.info

The Earth News is an independent English-language daily published Website from all around the World News

Browse by Category

  • Business (20,132)
  • Ecology (888)
  • Economy (910)
  • Entertainment (21,781)
  • General (17,834)
  • Health (9,951)
  • Lifestyle (923)
  • News (22,149)
  • People (911)
  • Politics (920)
  • Science (16,121)
  • Sports (21,410)
  • Technology (15,890)
  • World (893)

Recent News

The Bearded Vulture as an accumulator of historical remains: Insights for future ecological and biocultural studies – ESA Journals

The Bearded Vulture: Unlocking Secrets of Historical Remains for Future Ecological and Biocultural Research

October 27, 2025
«The subdued earth»: the history of thinking about science and faith. – omnesmag.com

The Subdued Earth: Unraveling the Intricate Relationship Between Science and Faith

October 27, 2025
  • About
  • Advertise
  • Privacy & Policy
  • Contact

© 2023 earth-news.info

No Result
View All Result

© 2023 earth-news.info

No Result
View All Result

© 2023 earth-news.info

Go to mobile version