Optimization & learning for design agents
Simulated annealing and reinforcement learning modules for reproducible engineering design research.
Overview
Design Research Agents is a modular Python framework for studying AI agents in engineering design, with common interfaces for model calls, tools, and multistep workflows. I implemented a simulated annealing search pattern and developed an episodic reinforcement learning pattern with the team. My work connected these algorithms to the framework’s existing workflows, configuration, and result reporting.
Simulated annealing
Design search can require exploring options that temporarily score worse before finding a better solution. I implemented simulated annealing with temperature-dependent acceptance of proposed designs. Higher temperatures allow more exploration, while cooling makes the search more selective.
I separated the search loop from the design problem so researchers could supply their own objective, constraints, and local changes. The module supports minimizing or maximizing an objective, supplied or generated starting states, and optional checks on the design representation.
I added adaptive cooling based on objective history. In a later API review, I made temperature schedules directly importable and added schedule settings and objective history to the results. I also fixed the output to preserve generated starting designs, making runs easier to inspect and analyze.
Reinforcement learning
We scoped the learning pattern around complete episodes because design tasks may only yield meaningful feedback after a finished design. I connected a user-defined environment to a policy that selects actions and learns from episode rewards.
I implemented an epsilon-greedy policy that balances random exploration with selecting the action with the highest estimated return. My initial version learned one value per discrete action from discounted episode rewards.
I defined the policy interface and implemented a workflow with episode and step limits, recording actions, rewards, and policy parameters. During iteration, I fixed state capture to preserve the history when an environment modifies a design in place.
I reviewed how repeated state-action pairs affected learning and changed the update to first-visit Monte Carlo. Only the return following a pair’s first occurrence in an episode updates its estimate.
Results
My code adds configurable simulated annealing and episodic reinforcement learning to the open-source library. Shared workflow and result contracts let researchers supply their own objectives, environments, and policies while retaining run records for analysis.