Researchers behind DS-STAR reported ablation studies to test how its individual components affect performance and to examine the impact of refinement rounds on the iterations needed to generate a sufficient plan.
The Data File Analyzer was described as essential for high performance. Without the descriptions it generates, listed as Variant 1, DS-STAR’s accuracy on difficult tasks in the DABStep benchmark dropped to 26.98%, highlighting the role of richer data context in planning and implementation.
The Router also proved important. When it was removed, in Variant 2, DS-STAR only added new steps sequentially, which led to worse performance on both easy and hard tasks. The results suggested that correcting mistakes in a plan was more effective than continuing to add potentially flawed steps.
The study also tested DS-STAR’s adaptability across language models by using GPT-5 as the base model. The GPT-5 version produced promising results on the DABStep benchmark, indicating generalizability.
According to the findings, DS-STAR with GPT-5 performed better on easy tasks, while the Gemini-2.5-Pro version performed better on hard tasks.
Source: research.google.
Companies can share verified announcements through Newz9’s international press release submission page.

