Skip to content
AGI4EngineeringPublic

About

A benchmark of 101 structured engineering design tasks across multiple domains.

Resources

Stars

14 stars

Watchers

1 watching

Forks

Repository files navigation

EngDesign

EngDesign is a benchmark of 101 structured engineering‑design tasks spanning multiple domains. This repository supports our NeurIPS Datasets & Benchmarks track submission:
“Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs.”

Among the 101 tasks in EngDesign, 48 require domain-specific scientific software such as MATLAB or Cadence for evaluation, which may not be excuable for all machines. The remaining 53 tasks are fully open-source and can be evaluated using manually authored scripts. To facilitate broader community adoption without licensing constraints, we have consolidated these 53 tasks into a subset called EngDesign-Open. Since the remaining 48 tasks depend on proprietary software, we currently provide run commands only for the open-source subset.


📂 Repository Layout

├── tasks/                 # 101 individual task folders
│   ├── <task_id>/        # e.g. XG_01
│   │   ├── LLM_prompt.txt      # Prompt presented to the LLM
│   │   ├── output_structure.py # Defines the expected JSON/Python output schema via instructor
│   │   ├── evaluate.py         # Runs simulations & computes evaluation results
│   │   ├── images/             # (Optional) Input images for multimodal tasks
│   │   └── logs/               # Our evaluation logs
│   └── ...
├── EngDesign-Open/
│   ├── <task_id>/
│   └── ...
├── iterative_result/      # Logs from iterative design runs with GPT‑4o, o1, o3, o4‑mini
└── evaluation/            # Driver scripts & helpers for running the benchmark
    ├── eval_openai_llm.py
    └── eval_openai_llm_new.py

🚀 How to Run All Open Source Tasks (EngDesign-Open)

EngDesign-Open contains all 53 open source tasks. You can run them by following these steps:

Setup Instructions

1. Install and Log in to Docker

  • Register at hub.docker.com and verify your email.
  • Download and install Docker Desktop on your machine: Download Docker Desktop
  • Launch Docker Desktop and log in to your account.
  • Make sure Docker Desktop has access to your drive (check settings).

2. Replace API Keys

Replace the top of evaluation/eval_openai_llm_new.py with your actual OpenAI API keys before building the container.

3. Authenticate via CLI

In a terminal, run:

docker login -u your_dockerhub_username

4. Build the Docker Image

Run the following command in the root directory of this project:

docker build -t engdesign-sim .

5. Start a Docker Container

Mount your local project directory and start a bash session in the container:

docker run -it --rm -v the_actual_full_path_to_your_local_project_directory --entrypoint bash engdesign-sim

6. Run the Benchmark Tasks

Once inside the container (you'll see a prompt like root@xxxxxxxxxxxx:/app#), run one of the following:

(1) Run All Tasks
xvfb-run -a -e /dev/stdout --server-args="-screen 0 1024x768x24" \
python3 evaluation/eval_openai_llm_new.py \
--task_dir ./EngDesign-Open --model gpt-4o --k 1
(2) Run Specific Tasks
xvfb-run -a -e /dev/stdout --server-args="-screen 0 1024x768x24" \
python3 evaluation/eval_openai_llm_new.py \
--task_dir ./EngDesign-Open --task_list AB_01 AB_02 --model gpt-4o --k 1
Parameter Description
Parameter Description
--task_dir Directory containing the task folders
--task_list (Optional) Names of specific tasks to run. If not set, all tasks will run
--model Model to use, e.g., gpt-4o
--k Number of repetitions per task

7. Exit the Container

Type exit to quit the container shell.

Optional Cleanup

Remove the image if needed:

docker image rm engdesign-sim

About

A benchmark of 101 structured engineering design tasks across multiple domains.

Resources

Stars

14 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages