AI for material selection

An evaluation pipeline for benchmarking AI material selection against expert judgment.

Overview

Material selection requires balancing performance, manufacturing, cost, and context. At Carnegie Mellon’s Design Research Collective, I investigated how model size, prompting, and access to search affect whether AI recommendations align with expert judgment.

I built the experiment pipeline, ran the studies, and analyzed the outputs as first author on two papers with Daniele Grandi, Allin Groom, and Christopher McComb.

Heatmap of mean absolute error between expert survey responses and Qwen3 ratings across five model sizes and five prompting methods
Fig. 1 Mean absolute error between Qwen3 ratings and expert responses across model sizes and prompting methods, where lower values indicate closer alignment

Research approach

We used an existing expert survey dataset, MSEval, with suitability ratings from 138 professionals across four design applications, four criteria, and nine material categories. I used the same 0–10 rating task for the models, comparing them with expert responses rather than assuming one correct answer.

We compared a search-equipped agent with four standalone prompting approaches across the Qwen 2.5 model family, keeping the design questions fixed. The prompts asked directly, supplied examples, separated explanation from scoring, or requested all nine material ratings together.

Agent implementation

I built the Python experiment pipeline by adapting generation and evaluation scripts from the earlier material selection study. Hugging Face’s ReAct code agent coordinates reasoning with Python tool calls, while llama-cpp-python runs the models locally. I added a Wikipedia search tool and a system prompt with examples outside the survey cases so the agent could inspect summaries before scoring.

Reviewing execution logs, I added checks for successful final answers, errors, and the rating range, with up to five attempts before omitting a failed response. Logs recorded searches, retries, and token usage to assess reliability and computational effort alongside alignment.

Expert comparison

I compared each score with matching expert responses using mean absolute error and z-scores, which express deviation relative to the expert distribution. Because signed deviations can cancel when averaged, I interpreted both metrics together. I used regression to examine model size, prompting, and their interaction.

I also compared search-query wording using text embeddings, numerical representations of sentence meaning. Removing design, criterion, and material names helped isolate query patterns beyond their subject matter.

Iteration

The initial study showed that increasing model size did not consistently improve alignment or execution. Extensive execution failures led us to exclude the 14B agent from most analyses, while standalone prompts achieved lower mean absolute error than the agent.

Qwen3 introduced built-in reasoning before the final answer, motivating us to repeat the comparison and examine whether reasoning changed the roles of prompting, size, and tool use. I extended the pipeline to Qwen3 and added reasoning diagnostics while retaining the benchmark. I also added R² checks to assess how well model scores predicted expert mean ratings.

Results

In the follow-up, reasoning increased token use, parallel prompting improved alignment compared with the earlier models, and agentic performance declined. The experiments exposed trade-offs between expert alignment, execution reliability, and computational effort.

These findings are specific to the tested models and benchmark. Wikipedia summaries, broad material categories, and the initial study’s single-run evaluation limit conclusions about practical engineering decisions.

Tools

PythonHugging Face