Introducing LLMTEXT, an open source toolkit for the llms.txt standard | Parallel
Introducing Parallel Search Turbo: the fastest and lowest-cost search for AI. Learn more.Learn more.
HumanMachine
October 30, 2025
\# Introducing LLMTEXT, an open source toolkit for the llms.txt standard
TL;DR: We're launching LLMTEXT, an open source toolkit that helps developers create, validate, and use llms.txt files—making any website instantly accessible to AI agents through standardized markdown documentation and MCP servers.
Tags:Product
Reading time: 6 min
\## The web wasn't built for AI agents.
Wikipedia recently reportedrecently reported an 8% decline in human visitors, and AI has overtaken humansAI has overtaken humans as the primary user of the web this year. As LLMs increasingly become the primary way people interact with online information, websites face a critical challenge: how do you serve both human visitors and AI agents effectively?
Today, we’re proud to support the launch of LLMTEXTLLMTEXT, an open source toolkit by Jan WilmakeJan Wilmake to help grow the llms.txt standard. With these tools, developers can more easily create llms.txt files for their websites, check their website’s existing llms.txt for validity, or turn any existing llms.txt into a dedicated MCP (Model Context Protocol) server.
\## What is LLMTEXT, and why did we make it?
Jeremy HowardJeremy Howard introduced the llms.txt standardllms.txt standard to make websites more friendly for large-language models (LLMs) by giving them access to Markdown files that contain the site’s most important text, explicitly excluding distracting elements that otherwise fill up their context windows. The spec has already been adopted by companies like Anthropic, Cloudflare, Docker, HubSpot, and many others.
At Parallel, we believe that AIs will soon be the primary users of the web, which is why we support initiatives like llms.txt that introduce new standards and frameworks to embrace that future.
\## What tools are available on LLMTEXT?
The three LLMTEXT tools we're releasing today serve two purposes. First, the **llms.txt MCP ** helps developers use projects without hallucination by getting a dedicated MCP for every library or API they use that supports llms.txt. On the other hand, the **Check tool** and **Create tool** aid websites to serve their users the best possible experience. Let's dive into each tool.
\## llms.txt MCP tool
The **llms.txt MCP ** turns any public llms.txt into a dedicated MCP server. You can think of this like Context7, but instead of one MCP server for all docs, it’s a narrowly-focused MCP server for a website, making it easier to get the right context for products you use often. It also works fundamentally differently: Where Context7 uses vector search to determine what’s relevant, the **llms.txt MCP ** leverages the reasoning of the LLM based on the llms.txt overview to decide which documents to ingest into the context window.
An LLMTEXT MCP being used in Cursor
Many developers have already been using llms.txt or the linked markdown files by manually copying them into their context window, but the llms.txt MCP smoothens this process by having one-click installation and providing the LLM with clear instructions on how to ingest the right context when needed. The MCP exposes two tools:
- **get - ** explains what this llms.txt is about, will first retrieve the llms.txt itself, then be used again to retrieve multiple relevant contexts for any task
- **leaderboard - ** shows the most active users of the MCP and other insights
The **llms.txt MCP ** can be installed for any llms.txt that follows the standard.
\## Check tool
When building the llms.txt MCP and trying it out on some of the available MCPssome of the available MCPs, many of them turned out to be incorrect according to the llms.txt prescribed format for various reasons.
An example of validation for Docker's llms.txt
The goal of your llms.txt should be to give LLMs the best possible overview: a table of contents to determine where to look for the right information. Or, as our Co-founder and Head of Product, TraversTravers, puts it, the goal is to retrieve the tokens that agents need to answer or make the next best decision in a loop. This means that you should have clear, distinct titles and descriptions of pages, and the individual results of pages shouldn't be too long.
To incentivize companies to fix their llms.txt and to provide only the best quality to our users, LLMTEXTLLMTEXT only allows installing MCP servers that adhere to the spec fully. Here are the most common mistakes we found in llms.txt files, with examples from popular websites:
\### Document Size
To get the most out of llms.txt, documents should be token-efficient. For example, https://developers.cloudflare.com/llms.txthttps://developers.cloudflare.com/llms.txt is 36,000 tokens for just the table of contents, creating a very large minimum amount of tokens.
Another example is https://docs.cursor.com/llms.txthttps://docs.cursor.com/llms.txt, which serves links to several languages. This isn't succinct and creates unnecessary overhead to an LLM that knows most languages.
To make token usage efficient when wading through context, it's best if the llms.txt itself is not bigger than the pages being linked to. If it is, it becomes a significant addition to the context window every time you want to retrieve a piece of information.
Another example is https://supabase.com/llms.txthttps://supabase.com/llms.txt, where the first document linked contains approximately 800,000 tokens, which is far too large for most LLMs to process. As a first rule, we recommend keeping both llms.txt and all linked documents under 10,000 tokens.
\### Incorrect content-type
The llms.txt itself, as well as the links it refers to, must lead to a text/markdown or text/plain response. This is the most common mistake in llms.txt files today.
For example, https://www.bitcoin.com/llms.txthttps://www.bitcoin.com/llms.txt and https://docs.docker.com/llms.txthttps://docs.docker.com/llms.txt both return HTML for every document linked to, and while listed in some registries, https://elevenlabs.io/llms.txthttps://elevenlabs.io/llms.txt responds with an HTML document.
In many cases, the content-type is text/plain or text/markdown, yet it can't be parsed according to the spec. For example, https://cursor.com/llms.txthttps://cursor.com/llms.txt just lists raw URLs without markdown link format, https://console.groq.com/llms.txthttps://console.groq.com/llms.txt does not present its links in an h2 markdown section (##), and https://lmstudio.ai/llms.txthttps://lmstudio.ai/llms.txt returns all documents directly, concatenated.
\### Not served at the root
Many companies ended up not serving their llms.txt at the root. For example, https://www.mintlify.com/docs/llms.txthttps://www.mintlify.com/docs/llms.txt and https://nextjs.org/docs/llms.txthttps://nextjs.org/docs/llms.txt are not hosted at the root, making it hard to find programmatically.
\## Create tool
Most websites aren't adapted to the AI internet yet, and instead serve as HTML content intended for humans. Most CMS systems don't support the creation of a Markdown versions either. There are several llms.txt generators (hosted as well as libraries) found on the internet, but many are specific to a certain framework. And many tools to create llms.txt don’t actually follow the llms.txt spec.
For example, some tools just create the llms.txt file itself, but don't refer to plain text or markdown variants of the pages.
Parallel's llms.txt, created with the create tool from LLMTEXT
The extract-from-sitemap tool is a framework-agnostic way to generate an llms.txt from multiple sources. It scrapes all needed pages and turns them into markdown, powered by the new Parallel Extract APIParallel Extract API. We used this library to create our own llms.txtour own llms.txt, which is also available through this repothis repo as reference, and installable as MCPinstallable as MCP for those building with Parallel's APIs.
\## Plans for LLMTEXT
This is just the beginning. We hope that the llms.txt standard thrives and evolves into a more valuable standard with many use-cases. We've already started improving the tooling and adding utilities, and hope to see the open source community contribute as well.
\## About Jan Wilmake
Jan has been an active OSS developer building dev tools in the AI context management space. His work includes uithub.comuithub.com, a context ingestion tool for GitHub, and the openapi-mcp-serveropenapi-mcp-server, which allows ingesting the full API specification of the operations you’re interested in, following a very similar pattern to how the llms.txt MCP works.
\## About Parallel Web Systems
Parallel develops critical web search infrastructure for AI. Our suite of web search and agent APIs is built on a rapidly growing proprietary index of the global internet. These solutions transform human tasks that previously took weeks into agentic tasks that now take just minutes.
Fortune 100 companies use Parallel’s search and agent APIs in insurance, finance, and retail, as well as AI-first businesses like Clay, Starbridge, and Sourcegraph.
\## Ready to get started?
Sign up for free. No credit card required.
Try ParallelTry Parallel Contact salesContact sales
Are you an agent? Read this to onboard ParallelAre you an agent? Read this to onboard Parallel
By Parallel
October 30, 2025
\## Related Posts80
Jul 30, 2026\ - Building an always-on background agent to proactively support customers
Author: By Khushi Shelat
Jul 21, 2026\ - Introducing the Parallel Responses API
Author: By Parallel
Jul 20, 2026\ - Building a vendor intelligence system with Parallel
Author: By Sahith Jagarlamudi
Jul 16, 2026\ - Parallel and Google Cloud Announce Partnership for Agentic Web Search on Gemini Enterprise Agent Platform
Author: By Parallel
Jul 15, 2026\ - $5 in free Parallel credits, every month
Author: By Parallel
\ \ Jul 13, 2026\ \ - [Introducing Parallel Search Turbo](/content/blog/parallel-search-turbo/index.html) \ \ Author: By Parallel](/content/blog/parallel-search-turbo/index.html)
Jul 12, 2026\ - Building a realtime voice agent with GPT-Realtime-2.1 and Parallel Search Turbo
Author: By George Pickett
Jul 10, 2026\ - How Nooks cut web search costs 70.5% by switching to Parallel
Author: By Parallel
Jul 8, 2026\ - How Build created live geofenced alerts powered by Parallel for institutional real estate
Author: By Parallel
Jun 9, 2026\ - OpenClaw now has free, LLM-optimized web search by default powered by Parallel
Author: By Parallel
Jun 5, 2026\ - Introducing real-time Entity Search
Author: By Parallel
Jun 3, 2026\ - How we enrich & triage inbound leads using the Parallel Task API
Author: By Khushi Shelat
May 20, 2026\ - How AirOps creates citation-worthy content at scale, powered by Parallel
Author: By Parallel
May 18, 2026\ - Introducing Index by Parallel
Author: By Parallel
May 7, 2026\ - Parallel Monitor API: New processor tiers, snapshots and event streams, and Basis on every event
Author: By Parallel
May 4, 2026\ - How we built parallelmpp.dev
Author: By Son Do
Apr 29, 2026\ - How Actively's Per Account Agents use Parallel to turn the entire web into a proactive sales intelligence layer
Author: By Parallel
Apr 28, 2026\ - Parallel Raises at $2 Billion Valuation to Scale Web Infrastructure for Agents
Author: By Parallel
Apr 24, 2026\ - Building a free CLI agent with Pi, Ollama, Gemma 4, and Parallel
Author: By Matt Harris
Apr 23, 2026\ - Parallel Search is now free for agents via MCP
Author: By Parallel
Apr 21, 2026\ - Upgrades to the Parallel Search & Extract APIs
Author: By Parallel
Apr 20, 2026\ - How Finch is scaling plaintiff law with AI agents that research like associates
Author: By Parallel
Apr 8, 2026\ - Genpact and Parallel Web Systems Partner to Drive Tangible Efficiency from AI Systems
Author: By Parallel
Apr 8, 2026\ - How Genpact helps top US insurers cut contents claims processing times in half with Parallel
Author: By Parallel
Apr 7, 2026\ - A new deep research frontier on DeepSearchQA with the Task API Harness
Author: By Parallel
Mar 30, 2026\ - How Modal saves tens of thousands annually by building in-house GTM pipelines with Parallel
Author: By Parallel
Mar 25, 2026\ - How Opendoor uses Parallel as the enterprise grade web research layer powering its AI-native real estate operations
Author: By Parallel
Mar 19, 2026\ - Introducing stateful web research agents with multi-turn conversations
Author: By Parallel
Mar 18, 2026\ - Parallel is live on Tempo, now available natively to agents with the Machine Payments Protocol
Author: By Parallel
Mar 17, 2026\ - How Parallel helped Kepler build AI that finance professionals can actually trust
Author: By Parallel
Mar 10, 2026\ - Introducing the Parallel CLI
Author: By Parallel
Mar 4, 2026\ - How Profound helps brands win AI Search with high-quality web research and content creation powered by Parallel
Author: By Parallel
Mar 2, 2026\ - How Harvey is expanding legal AI internationally with Parallel
Author: By Parallel
Feb 23, 2026\ - How Tabstack by Mozilla enables agents to navigate the web with Parallel’s best-in-class web search
Author: By Parallel
Feb 4, 2026\ - Parallel Web Tools and Agents now available across Vercel AI Gateway, AI SDK, and Marketplace
Author: By Parallel
Jan 28, 2026\ - Authenticated page access for the Parallel Task API
Author: By Parallel
Jan 21, 2026\ - Introducing structured outputs for the Monitor API
Author: By Parallel
Jan 15, 2026\ - Introducing research models with Basis for the Parallel Chat API
Author: By Parallel
Jan 8, 2026\ - Build a real-time fact checker with Parallel and Cerebras
Author: By Parallel
Dec 17, 2025\ - Parallel Task API achieves state-of-the-art accuracy on DeepSearchQA
Author: By Parallel
Dec 16, 2025\ - Introducing Granular Basis for the Task API
Author: By Parallel
Dec 11, 2025\ - How Amp’s coding agents build better software with Parallel Search
Author: By Parallel
Dec 10, 2025\ - Latency improvements on the Parallel Task API
Author: By Parallel
Nov 20, 2025\ - Introducing Parallel Extract
Author: By Parallel
Nov 18, 2025\ - Introducing Parallel FindAll
Author: By Parallel
Nov 13, 2025\ - Introducing Parallel Monitor
Author: By Parallel
Nov 12, 2025\ - Parallel raises $100M Series A to build web infrastructure for agents
Author: By Parallel
Nov 11, 2025\ - How Macroscope reduced code review false positives with Parallel
Author: By Parallel
Nov 6, 2025\ - Introducing Parallel Search
Author: By Parallel
Nov 3, 2025\ - Parallel processors set new price-performance standard on SealQA benchmark
Author: By Parallel
Oct 23, 2025\ - How Starbridge powers public sector GTM with state-of-the-art web research
Author: By Parallel
Oct 22, 2025\ - Building a market research platform with Parallel Deep Research
Author: By Parallel
Oct 17, 2025\ - How Lindy brings state-of-the-art web research to automation flows
Author: By Parallel
Oct 16, 2025\ - Introducing the Parallel Task MCP Server
Author: By Parallel
Oct 9, 2025\ - Introducing the Core2x Processor for improved compute control on the Task API
Author: By Parallel
Oct 8, 2025\ - How Day AI merges private and public data for business intelligence
Author: By Parallel
Oct 7, 2025\ - Full Basis framework for all Task API Processors
Author: By Parallel
Oct 6, 2025\ - Building a real-time streaming task manager with Parallel
Author: By Parallel
Sep 30, 2025\ - How Gumloop built a new AI automation framework with web intelligence as a core node
Author: By Parallel
Sep 16, 2025\ - Introducing the TypeScript SDK
Author: By Parallel
Sep 12, 2025\ - Building a serverless competitive intelligence platform with MCP + Task API
Author: By Parallel
Sep 11, 2025\ - Introducing Parallel Deep Research reports
Author: By Parallel
Sep 9, 2025\ - A new pareto-frontier for Deep Research price-performance
Author: By Parallel
Sep 5, 2025\ - Building a Full-Stack Search Agent with Parallel and Cerebras
Author: By Parallel
Aug 21, 2025\ - Webhooks for the Parallel Task API
Author: By Parallel
Aug 14, 2025\ - Introducing Parallel: Web Search Infrastructure for AIs
Author: By Parallel
Aug 7, 2025\ - Introducing SSE for Task Runs
Author: By Parallel
Aug 5, 2025\ - A new line of advanced Processors: Ultra2x, Ultra4x, and Ultra8x
Author: By Parallel
Aug 4, 2025\ - Introducing Auto Mode for the Parallel Task API
Author: By Parallel
Jul 31, 2025\ - A state-of-the-art search API purpose-built for agents
Author: By Parallel
Jul 31, 2025\ - Parallel Search MCP Server in Devin
Author: By Parallel
Jul 28, 2025\ - Introducing Tool Calling via MCP Servers
Author: By Parallel
Jul 14, 2025\ - Introducing the Parallel Search MCP Server
Author: By Parallel
Jul 8, 2025\ - Introducing Source Policy
Author: By Parallel
Jul 2, 2025\ - The Parallel Task Group API
Author: By Parallel
Jun 17, 2025\ - State of the Art Deep Research APIs
Author: By Parallel
Jun 10, 2025\ - Parallel Search API is now available in alpha
Author: By Parallel
May 29, 2025\ - Introducing the Parallel Chat API
Author: By Parallel
May 16, 2025\ - Introducing Basis with Calibrated Confidences
Author: By Parallel
Apr 24, 2025\ - Introducing the Parallel Task API
Author: By Parallel