AI models only half right on industrial parts questions, vendor test shows

Reshape Automation’s industrial AI accuracy index tested AI models on questions distributors and OEMs get

Reshape Automation’s new Industrial AI Accuracy Index 2026 tested three artificial intelligence (AI) models, GPT-6 Astra, Claude and Gemini, on 100 questions drawn from what customers ask distributors and OEMs every day, covering parts from 14 industrial manufacturers, including Siemens, Festo, Rittal and ATI Industrial Automation. From memory alone, the models got 12% to 14% right.

With web search, the best scored 52%. About one in three wrong answers gave a specific part number or figure with no caveat. One of the 100 industrial parts questions in Reshape Automation’s report: A customer asks their distributor whether a Siemens communication module, part number 3RW5950-0CH00, works with a 3RW52 soft starter, and the distributor turns to AI to find out. Claude, with web search turned on, said yes all three times it was asked. The verified answer is no. The module belongs with the 3RW55 series, and an order placed on that answer ships a part that doesn’t fit the equipment it was bought for.

What the models got wrong

With web search, GPT-6 Astra scored 52% and Claude 40.5%. Without it, all three models landed between 12% and 14%. Even the best setup got 42 of the 100 questions wrong, with partial answers earning half credit.

About one in three wrong answers gave a specific code or figure with no caveat. Asked when a Festo actuator was discontinued, one model gave a different part number and date on each of three tries, and none was correct.

Each question went to five model setups, three times each, and in 66 of those 500 pairings the runs didn’t agree. Asked for a Rittal KX terminal box in 304 stainless, 200 x 200 x 80, Gemini gave a specific SKU and then said no such box exists in that range and then gave a different SKU.

Where a part number is built from the manufacturer’s ordering rules instead of printed in a catalog, as with ATI’s robot tool changers, GPT-6 Astra with web search scored 22.7% and Claude scored 0%. On ordinary part lookups, web search had raised scores by 51 points for GPT-6 Astra and 35 for Claude.

Why the models get parts questions wrong

“The AI products on the market today are mostly a general-purpose model with search or a document store bolted on, and they’re very good at a lot of work,” said Juan Aparicio, CEO and co-founder of Reshape Automation. “Parts questions are different. The answer is usually a relationship between three or four facts that live in different places, like a catalog, a datasheet or a cross-reference table, and some of it was never written down at all. If the model has to rebuild that relationship every time someone asks, it’ll sometimes get it right and sometimes not, and it’ll sound equally sure both times.”

“Distributors are where these questions land,” Aparicio said. “A customer asks whether a part fits, someone on the inside sales team has a few minutes to answer, and one wrong digit ships the wrong part. A tool that’s right half the time doesn’t save that person any work, because they still have to check every answer.”

Reshape agents

Reshape Automation builds AI agents that take on the technical and commercial work of industrial sales and support. Its ReshapeX agents select and configure products, cross-reference parts and full bills of materials, prepare quotes, track orders, guide field troubleshooting and surface account insights for distributors, OEMs, system integrators and manufacturers. They work inside the customer’s ERP and CRM and run on websites, Outlook, Microsoft Teams, WhatsApp and voice.

Reshape’s engineers wrote up the 100 questions from what customers ask the distributors and OEMs Reshape works with, and no third party reviewed them. The answer key comes from ReshapeX’s own answers, verified against manufacturer documentation, so ReshapeX is not scored on equal terms with the models it tested. Every question is published in the report so anyone can rerun the test.

Each model ran through its provider’s API with a pinned model ID (gpt-6-astra, claude-fable-5-1, gemini-3.8-flash): 1,500 calls on September 24. Gemini ran without web search because its provider’s terms require written permission to benchmark it with grounding. A model drafted each grade and Reshape’s engineers made the final call. The report lists 13 limitations.

The report prints all 100 questions and an SHA-256 fingerprint of the answer key. The key itself is withheld to keep the questions usable as a test and is available on request.

Reshape Automation’s ReshapeX family of AI agents build valid part numbers from specs, find replacements and accessories, and cross-reference and price a bill of materials hundreds of lines long.

Item, the German maker of aluminum building-kit systems, runs a ReshapeX agent on its U.S. and Mexico websites in English and Spanish. The agent answers from Item’s own product data across more than 4,500 mutually compatible components, an,d when a question goes past what it can verify, it hands the customer to an Item expert with the whole conversation attached.

“Most AI on a website waits for you to ask and forgets you the moment you leave. The item agent is part of Item’s journey, not a box bolted onto it,” said Germain Dufossé, CEO of Item America.

Sign up for our eNewsletters
Get the latest news and updates