---
source_url: "https://www.wolfram.com/llm-benchmarking-project/"
title: Wolfram LLM Benchmarking Project
mirrored_at: 2026-08-30T03:02:21.039Z
host: www.wolfram.com
cited_in_42a: true
mirror_canonical: "https://index.42a.ai/www.wolfram.com/llm-benchmarking-project/index"
---

> **Original source:** https://www.wolfram.com/llm-benchmarking-project/

As major users and analyzers of large language model (LLM) technology, we've been continually tracking the performance of LLMs. This project involves releasing our ongoing results, initially for a specific well-characterized code generation task.

The task consists of going from English-language specifications to Wolfram Language code. The test cases are exercises from [Stephen Wolfram's _An Elementary Introduction to the Wolfram Language_](https://www.wolfram.com/language/elementary-introduction/3rd-ed/). These exercises have been done online by millions of humans, and we've developed effective tools for determining functional correctness of code, which we're now applying to LLMs.

This table and previous versions are available in computable form in the [Wolfram Data Repository](https://datarepository.wolframcloud.com/resources/LLMBenchmarks-Data/).

Find out how **Wolfram Language** can [enhance your LLM results](https://www.wolfram.com/artificial-intelligence/foundation-tool).

**For LLM developers**: [contact us](mailto:partner-program@wolfram.com) for the dataset and tools or to arrange for your LLM to be included.