Skip to content

Repository files navigation

LongCat-DeepResearch

Plan globally. Research independently. Synthesize coherently.

Project blog Paper PDF GitHub repository MIT license

English | 简体中文

Overview

This repository open-sources the LongCat-DeepResearch harness for open-ended research. It captures evolving research requirements in an executable ResearchSpec, supports independent section research and writing, and makes targeted revisions to the assembled report instead of repeatedly rewriting it. Users can run deep-research tasks with their own services by implementing llm, web_search, and web_fetch according to the backend contract. Paired with our LongCat-2.0-based model enhanced for deep research, the complete system outperforms the Deep Research offerings from ChatGPT, Claude, and Gemini on DeepResearchBench, DeepResearchBench II, and ResearchRubrics.

LongCat-DeepResearch benchmark overview

Research Harness

LongCat-DeepResearch harness

The harness follows three stages:

  1. Explore and plan. Multiple planners inspect the problem and early sources. A judge merges proposals, a critic identifies gaps, and a reviser produces the executable ResearchSpec.
  2. Research sections. Independent researchers investigate and write assigned sections in separate contexts while retaining the full specification.
  3. Assemble and edit. Completed sections are assembled directly. A global editor identifies ownership and consistency issues, and local editors apply section-scoped revisions.

Research Data Construction

Evidence-grounded research data construction

The evidence-grounded research-data construction pipeline builds research questions and task-specific rubrics from an independent CC-BY review article or a frozen multi-source brief, validates evidence support, searchability, leakage, and task quality, and collects and filters stage-specific harness trajectories from accepted queries.

This repository releases only the research harness. The data-construction pipeline, generated datasets, hidden rubrics, training trajectories, and training code are not included.

Quick Start

The harness itself uses only the Python standard library and supports Python 3.10 or newer. Install the repository in editable mode:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e .

To run it, provide one Python object with three methods:

  • llm sends an OpenAI Chat Completions/function-calling request to your model and returns the provider's response as a dictionary.
  • web_search receives one query and returns normalized search results.
  • web_fetch receives one or more URLs and returns normalized readable pages in the same order.

No hosted backend is bundled because model, search, and page-fetching services normally require account-specific credentials. Create local_backend.py in the repository root. It is already ignored by Git, so credentials and provider-specific code remain local:

class Backend:
    def llm(self, request: dict) -> dict:
        # request: {"model", "messages", "stream", "max_tokens", ...}
        # return: {"choices": [{"message": {...}, "finish_reason": "..."}]}
        raise NotImplementedError

    def web_search(self, request: dict) -> dict:
        # request: {"query": str, "top_k": int}
        # return: {"results": [{"title": str, "url": str,
        #                       "snippet": str, "published_at": str | None}]}
        raise NotImplementedError

    def web_fetch(self, request: dict) -> dict:
        # request: {"urls": [str, ...]}
        # return: {"pages": [{"url": str, "title": str,
        #                     "text": str, "error": str | None}, ...]}
        raise NotImplementedError


def create_backend():
    return Backend()

Set the module and model name, then verify the three configured services. The smoke command makes one small model request, one search request, and one page-fetch request; it does not generate a research report:

export DR_BACKEND_MODULE=local_backend
export DR_MODEL=your-model-name

scripts/run_smoke.sh

Run the complete research harness separately:

export DR_BACKEND_MODULE=local_backend
export DR_MODEL=your-model-name

printf '%s\n' 'What question should the system research?' > question.txt
scripts/run_standard.sh question.txt runs

The final report is written to runs/<question_hash>/report_final.md. Intermediate planning, section research, and editing artifacts remain under the same run directory for inspection.

For library use, inject the backend directly:

from longcat_deepresearch import LongCatDeepResearch

result = LongCatDeepResearch(backend=my_backend).run(
    "What question should the system research?",
    run_root="runs",
)
print(result.final_path)

Backend Contract

The backend is a transport boundary: the harness decides when to call the model and tools, while your backend decides how to authenticate and talk to external services. The harness passes web_search and web_fetch as OpenAI-style function tools in llm requests. If the model returns a tool call, the harness executes the corresponding backend method and sends the tool result back to the model in the next turn.

The public interface is defined in longcat_deepresearch/backend.py. Every method accepts one Python dictionary and returns one Python dictionary:

Method Request supplied by the harness Required response
llm model, messages, stream=false, max_tokens, and optional tools/tool_choice An OpenAI Chat Completions-compatible object containing choices[0].message
web_search {"query": "...", "top_k": 8}; top_k is between 1 and 10 {"results": [...]} with normalized title, URL, snippet, and optional publication date
web_fetch {"urls": ["https://..."]} with 1–10 HTTP(S) URLs {"pages": [...]} with exactly one page per requested URL, in request order

For a terminal model answer, llm returns:

{
  "choices": [
    {
      "message": {"role": "assistant", "content": "The answer or report text."},
      "finish_reason": "stop"
    }
  ]
}

For a model-requested tool call, message.content may be null. The function arguments value must be a JSON-encoded string, not a nested object:

{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_1",
            "type": "function",
            "function": {
              "name": "web_search",
              "arguments": "{\"query\":\"example\",\"top_k\":5}"
            }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ]
}

A normalized search response looks like this:

{
  "results": [
    {
      "title": "Source title",
      "url": "https://example.org/source",
      "snippet": "Short search-result summary.",
      "published_at": "2026-01-01"
    }
  ]
}

A normalized fetch response looks like this:

{
  "pages": [
    {
      "url": "https://example.org/source",
      "title": "Source title",
      "text": "Readable page text.",
      "error": null
    }
  ]
}

See the complete backend interface for field-level rules and a multi-URL fetch example. The backend owns authentication, networking, redirects, timeouts, retries, rate limits, provider-specific fields, response-size limits, thread safety, and private/link-local network blocking.

Repository Structure

.
├── assets/                 # README figures and visual assets
├── docs/                   # Detailed backend contract
├── longcat_deepresearch/   # Installable Python package and CLI
├── scripts/                # CLI and release checks
├── technical_report/       # Latest technical report PDF
├── tests/                  # Offline interface and end-to-end tests
└── pyproject.toml          # Package metadata and development tooling

Testing

All tests use fake backends and make no external requests.

python -m pip install -e ".[dev]"
python3 -m unittest discover -s tests -v
ruff check . --exclude .venv
python3 scripts/open_source_preflight.py

Limitations and Responsible Use

Generated reports may contain unsupported claims, incorrect citations, unsafe recommendations, or omissions introduced during editing. Operators must review outputs, respect website terms and robots policies, protect personal information, and obtain authorization for every model, dataset, and service connected to the harness.

Backend implementations must treat retrieved content as untrusted input, prevent access to private and link-local network destinations, protect credentials, and avoid logging authorization headers or private prompts.

Do not disclose vulnerabilities in public issues. Use the repository's private security-reporting channel and include the affected version, reproduction steps, and expected impact.

Citation

If you find this project useful, please cite:

@techreport{longcatdeepresearch2026,
  title  = {LongCat-DeepResearch Technical Report},
  author = {
    Meituan LongCat Team and He Zhu and Yue Xu and
    Xunliang Cai and Yan Chen and Fan Yang and
    Lingchuan Liu and others
  },
  year   = {2026}
}

License

Copyright 2026 LongCat.

This project is released under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

20 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages