Skip to content
the mythical llama-month

10X coders beware: Meta’s new AI model boosts coding and debugging for free

New “Code Llama” coding model is free for research and commercial use.

Benj Edwards – | 117
Story text

Meta is adding another Llama to its herd—and this one knows how to code. On Thursday, Meta unveiled “Code Llama,” a new large language model (LLM) based on Llama 2 that is designed to assist programmers by generating and debugging code. It aims to make software development more efficient and accessible, and it’s free for commercial and research use.

Much like ChatGPT and GitHub Copilot Chat, you can ask Code Llama to write code using high-level instructions, such as “Write me a function that outputs the Fibonacci sequence.” Or it can assist with debugging if you provide a sample of problematic code and ask for corrections.

As an extension of Llama 2 (released in July), Code Llama builds off of weights-available LLMs Meta has been developing since February. Code Llama has been specifically trained on source code data sets and can operate on various programming languages, including Python, Java, C++,  PHP, TypeScript, C#, Bash scripting, and more.

Notably, Code Llama can handle up to 100,000 tokens (word fragments) of context, which means it can evaluate long programs. To compare, ChatGPT typically only works with around 4,000-8,000 tokens, though longer context models are available through OpenAI’s API. As Meta explains in its more technical write-up:

Aside from being a prerequisite for generating longer programs, having longer input sequences unlocks exciting new use cases for a code LLM. For example, users can provide the model with more context from their codebase to make the generations more relevant. It also helps in debugging scenarios in larger codebases, where staying on top of all code related to a concrete issue can be challenging for developers. When developers are faced with debugging a large chunk of code they can pass the entire length of the code into the model.

Meta’s Code Llama comes in three sizes: 7, 13, and 34 billion parameter versions. Parameters are numerical elements of the neural network that get adjusted during the training process (before release). More parameters generally mean greater complexity and higher capability for nuanced tasks, but they also require more computational power to operate.

A demonstration of Code Llama provided by Meta.
A demonstration of Code Llama provided by Meta. Credit: Meta

The different parameter sizes offer trade-offs between speed and performance. While the 34B model is expected to provide more accurate coding assistance, it is slower and requires more memory and GPU power to run. By contrast, the 7B and 13B models are faster and more suitable for tasks requiring low latency, like real-time code completion, and can run on a single consumer-level GPU.

Meta has also released two specialized variations: Code Llama-Python, and Code Llama-Instruct. The Python variant is optimized specifically for Python programming (“fine-tuned on 100B tokens of Python code”), which is an important language in the AI community. Code Llama-Instruct, on the other hand, is tailored to better interpret user intent when provided with natural language prompts.

Additionally, Meta says the 7B and 13B base and instruct models have been trained with “fill-in-the-middle” (FIM) capability, which allows them to insert code into existing code, which helps with code completion.

License and data set

Code Llama is available with the same license as Llama 2, which provides weights (the trained neural network files required to run the model on your machine) and allows research and commercial use, but with some restrictions laid out in an acceptable use policy.

Meta has repeatedly stated its preference for an open approach to AI, although its approach has received criticism for not being fully “open source” in compliance with the Open Source Initiative. Still, what Meta provides and allows with its license is far more open than OpenAI, which does not make the weights or code for its state-of-the-art language models available.

Meta has not revealed the exact source of its training data for Code Llama (saying it’s based largely on a “near-deduplicated dataset of publicly available code”), but some suspect that content scraped from the StackOverflow website may be one source. On X, Hugging Face data scientist Leandro von Werra shared a potentially hallucinated discussion about a programming function that included two real StackOverflow user names.

In the Code Llama research paper, Meta says, “We also source 8% of our samples data from natural language datasets related to code. This dataset contains many discussions about code and code snippets included in natural language questions or answers.”

Still, von Werra would like to see specifics cited in the future. “It would be great for reproducibility and sharing knowledge with the research community to disclose what data sources were used during training,” von Werra wrote. “Even more importantly it would be great to acknowledge that these communities contributed to the success of the resulting models.”

How well does it perform?

To evaluate the performance of Code Llama, Meta used the HumanEval metric introduced by OpenAI in 2021. The HumanEval consists of 164 handcrafted programming problems aimed at testing the capabilities of code generation models.

While Meta has positioned Code Llama as a noteworthy competitor in AI-driven coding assistance, a nuanced examination of its performance metrics reveals a more complex picture. According to HumanEval benchmark tests, Code Llama’s Python-specialized 34B model scored 53.7 percent, the highest among weights-available models. Unsurprisingly, this performance is eclipsed by GPT-4, which achieved 67 percent on the same benchmark but packs a much higher parameter count—and potentially multiple models under the hood.

The foundational models of Code Llama—featuring 7B, 13B, and 34B parameters—scored 33.5, 36, and 48.8 percent, respectively, on HumanEval. These numbers are markedly lower than GPT-4’s score. The fine-tuned Instruct variants of Code Llama also failed to outperform GPT-4, with the Instruct 34B model scoring 41.5 percent.

A chart of Code Llama performance metrics provided by Meta.
A chart of Code Llama performance metrics provided by Meta.
A chart of Code Llama performance metrics provided by Meta. Credit: Meta

However, one key advantage that Code Llama has over GPT-4 is its accessibility. As previously mentioned, Code Llama is free and can run on your local machine, which has distinct privacy advantages, especially when working with proprietary code. GPT-4, on the other hand, requires a subscription (or paid access per token through the API)—and all the data it processes gets sent to OpenAI through the cloud, which can be a no-no for some businesses even if OpenAI promises not to use your data for nefarious purposes.

Interestingly, the Code Llama research paper also mentions an unreleased model called “Unnatural Code Llama” trained on LLM-generated examples that has been turning heads on social media because it achieves 62.2 percent on the HumanEval, which is very close to GPT-4’s 67 percent benchmark result. But Meta is keeping that model close to its vest for now.

To use Code Llama at the moment, you’ll need some technical experience setting up software, but it’s likely that people will soon create more user-friendly interfaces. AI researcher Simon Willison told Ars, “I’m excited to see if someone can build a good locally hosted copilot on top of it with a whole lot of extra work. The non-instruct models look like they are intended to support that use-case.”

While Code Llama appears to offer potential benefits for increasing productivity in coding, its effectiveness in real-world scenarios remains to be seen. Like other AI models, it could struggle with nuanced tasks or inadvertently introduce errors and security vulnerabilities. Meta hopes that Code Llama will pave the way for more specialized tools built upon Llama 2, but the full range of potential uses and the model’s limitations in a professional coding environment have yet to be fully explored.

You can request access to the Code Llama weights by filling out a form on the Meta website. The inference code required for running Code Llama is available on GitHub.

Photo of Benj Edwards
Benj Edwards Senior AI Reporter
Benj Edwards was a reporter at Ars Technica covering artificial intelligence and technology history.
117 Comments