Connecting your business to the machine: catalog, schema, and crawlers
It's the least glamorous, most lucrative work in GEO: making your site legible to the machine. You open the right robots in robots.txt (without this, you're not even in the game), you mark your brand and questions in JSON-LD with short, copyable answers, you understand what llms.txt is (and isn't), and you organize your product's Golden Record. It mirrors the 'connect to the data' beat from other tracks, only here the data you're connecting is your own business, and the reader is the AI that recommends you.
You spent months polishing the site, the brand's story, the product photos. Beautiful. Then AI comes along to recommend someone in your category and... skips you. Not because you're worse. Because your site isn't written in a way the machine can read, copy, and cite with confidence. It's like having the best shop on the street with the door locked on exactly the day a new customer walked by. This lesson is the key: opening the right door (the robots), putting up the sign the machine reads (schema), and organizing the storefront so it can describe your product without inventing things.
Think of a controller who builds a flawless balance sheet, but keeps it all in a scanned PDF no system can read. The number is right, just illegible to the machine. In marketing it's the same: your brand can be great, but if the site has no structured sheet (schema), AI can't extract "who it is, what it sells, what the price is" with confidence, so it recommends the competitor whose sheet is clean and complete. Connecting the catalog to the machine is the equivalent of handing over already-reconciled data instead of forcing AI to guess.
You wouldn't file a petition without precisely qualifying the parties, right? Full name, tax ID, address, everything in the right field so it doesn't get thrown out. Schema is qualifying your business's parties for AI: here's the brand's official name, here's what it does, here are the frequently asked questions with the exact answer. Without this, AI "qualifies" you on its own, stitching together loose pieces of the site, and sometimes gets it wrong. Marking up Organization and FAQ keeps the machine from filling the gap with invention.
You've been running ads, posting, doing SEO for years to get clicked. Fine. Except now a huge chunk of the decision happens inside a conversation with AI, where nobody clicks a link, they just read the answer. And to get into that answer the game is different: AI needs to be able to read your site, find your brand as a clear entity, copy an answer of yours straight into its text. This doesn't get solved with more creative work. It gets solved in the technical basement: unblocked robots, marked-up schema, complete product sheet. Less glamour, more money.
Think of that job posting you believe is well written, but the candidate sends a résumé in a scanned PDF your ATS can't read. The talent might be great, just illegible to the machine, so they never even make the screening. In marketing it's the same: your brand can be excellent, but if the site has no structured sheet (schema), AI can't extract "who it is, what it sells, what the price is" with confidence, and recommends the competitor whose sheet is clean. Connecting the catalog to the machine is like handing over the already-parsed profile, instead of forcing the system to guess who that candidate is.
You wouldn't throw a feature onto the roadmap without a PRD, right? Without acceptance criteria, without a metric, the engineering team "interprets" what you meant and sometimes builds something else. Schema is your business's PRD for AI: here's the brand's official name, here's what it does, here are the frequently asked questions with the exact answer, all in a structured field. Without this, AI discovers you on its own, stitches together loose signals from the site, and prioritizes the competitor it understood better. Marking up Organization and FAQ keeps the machine from doing discovery in the dark.
Think of that scorching-hot lead that lands in the CRM, but with half the fields empty: no company, no title, no context. The team can't qualify it, the forecast gets skewed, and the deal slides to the salesperson who had the complete sheet. In marketing it's the same logic: if your site has no schema, AI can't read "who the brand is, what it sells, what the price is" with confidence, so it sends the recommendation to the competitor whose profile is clean. Connecting the catalog to the machine means handing over the already-enriched lead, instead of forcing AI to build the proposal by guessing.
Think of a shipment where the package label came out smudged: the product is right, but the barcode scanner doesn't match, so the order stalls and blows the SLA. In marketing it's the same: your brand can be great, but if the site has no structured sheet (schema), AI can't read "who it is, what it sells, what the price is" with confidence, and recommends the competitor whose sheet scans cleanly on the belt. Connecting the catalog to the machine is like standardizing the label so the scanner recognizes it on the first pass, instead of forcing the process to stop and guess what's inside the box.
You wouldn't close an audit with a poorly documented control, a vague evidence field, and no named owner, right? The auditor won't sign off, they'll flag the nonconformity. Schema is your business's documentation for AI: here's the brand's official name, here's what it does, here's the frequently asked question with the exact answer, all traceable. Without this, AI fills the gap on its own, stitching together loose pieces of the site, and sometimes gets the facts about you wrong. Marking up Organization and FAQ closes the control with clear evidence, instead of letting the machine infer what was missing.
Think of a service deployed without API documentation: the endpoint exists, but since no one described the contract, the client integrates by guessing and the call breaks in production. Schema is your legible contract for AI: here's the brand's official name, here's what it does, here's the FAQ with the exact answer, all in a structured field the machine consumes without parsing loose HTML. Without this, AI "infers" your brand from raw text, gets the facts wrong, and recommends the competitor whose schema is well defined. Connecting the catalog to the machine means publishing the contract, instead of leaving the integration to reverse engineering.
Think of a lovely flow in the prototype that, in user testing, breaks down because the button label is ambiguous and the person doesn't know what to click. The path exists, just illegible to whoever's walking through it. In marketing it's the same: your brand can be great, but if the site has no structured sheet (schema), AI can't read "who it is, what it sells, what the price is" clearly and recommends the competitor whose journey it fully understands. Connecting the catalog to the machine is like labeling every step of the flow with no ambiguity, instead of forcing whoever arrives to guess where to go.
You wouldn't walk a board deck onstage without a source for the number on the slide, right? Without the citation, the board doesn't trust the data and the decision stalls. Schema is your business's citable source for AI: here's the brand's official name, here's what it does, here's the frequently asked question with the exact answer, all in a structured field. Without this, AI builds the "briefing" on you by itself, stitching together loose pieces of the site, and sometimes gets the positioning wrong. Marking up Organization and FAQ means handing over the source ready-made, instead of letting the machine guess at the data.
Let me warn you upfront: this is the most "basement" lesson in the entire marketing track. No creative, no brilliant copy, no campaign. It's pipes, wiring, and signage. And that's exactly why it's so lucrative, because almost nobody does it, and whoever does shows up in AI's answer while the competitor, even with a better brand, stays invisible. Think with me: what's the point of AI wanting to recommend you if, when it goes to read your site, it finds a locked door or a storefront it can't describe? This is the lesson that opens the door and tidies up the storefront.
The core idea of this lesson. To be found and recommended by AI, before any pretty strategy, you need to make your business legible to the machine. There are four pieces, from most basic to most refined. First, unblock the right robots in robots.txt (GPTBot, PerplexityBot, ClaudeBot, Google-Extended), because if you block them, you're not even in the game. Second, mark up your brand (Organization) and your frequently asked questions (FAQ) in JSON-LD, with each answer written in 40 to 80 words, self-contained, copyable. Third, know llms.txt without kidding yourself: adoption around 10%, no official support from big tech, it's "hype to watch", not a silver bullet. Fourth, the product's Golden Record: the structured, near-complete sheet that gets AI to recommend you with confidence. It's this track's "connect to the data" beat, only the data you connect is your own business.
01The right robots: if you locked the door, you're already done here
Let's start with the silliest, most fatal one. There's a small text file at the root of every site, robots.txt, that says which robots can enter and read your content. It's existed for decades, always serving Google's robot. Now there are new robots, the AIs', and each has a name. The ones that matter to you today:
- GPTBot, OpenAI's crawler (the one that feeds ChatGPT).
- PerplexityBot, Perplexity's.
- ClaudeBot, Anthropic's (Claude's robot).
- Google-Extended, Google's separate control for its generative AI features.
Here's the trap that catches a lot of good businesses: in the wave of "protecting content from AI", many companies blocked these robots in robots.txt. It makes sense for a newspaper selling subscriptions. For you, who WANTS to be recommended, it's a silent shot in the foot. If AI can't read your site, it has no way of citing you. You removed yourself from the answer before the game even started, and you didn't even notice.
The rule is direct: if your goal is to be found and recommended, these robots need to be unblocked. It's not an advanced strategy, it's the prerequisite for everything that comes after. Before you think about schema, content, anything, open the door.
An honest caveat so you don't get confused: unblocking the robot doesn't mean opening everything to anyone without criteria. You still decide what's public (the product page, the FAQ, the institutional page) and what stays out (logged-in areas, customer data). The point is not to lock, out of carelessness or trend-following, exactly the pages you'd want AI to recommend. Fair enough?
02Schema: the sign the machine reads (and copies)
With the door open, comes the second piece. AI can even read the loose text on your page, but guessing "what's the brand's official name, what it does, what the exact answer to this question is" costs effort, and wherever there's effort there's error. Schema solves this. Schema is a small piece of code (called JSON-LD) you put on the page that says, in a structured, unambiguous way: this here is an Organization, its name is such-and-such, it does such-and-such. Or: this is a frequently asked question, and this is the exact answer.
You don't need to write that code by hand. Every serious site platform (and any dev, or AI itself) generates it in minutes. Your marketing work, the part nobody does for you, is deciding the CONTENT. And here lives the trick, worth the gold in this lesson:
For the board. Mark up two things first: Organization (your brand's sheet as a clear entity, with name, what it does, official links) and FAQ (the questions your customer actually asks, with the answer right below). And write each answer as a self-contained block of 40 to 80 words: starts by answering directly, doesn't depend on the previous paragraph, and can be copied whole into AI's answer without losing meaning. That's what gets AI to cite you verbatim, in your own words, instead of paraphrasing the competitor.
Why 40 to 80 words? Because AI builds its answer in short, complete pieces. A long paragraph, full of "as we saw above", it can't cleanly cut out. A block that answers the question from start to finish, about the size of a good long tweet, it copies whole. GEO experts call this extractable content: you write already thinking about being clipped out.
And a serious warning, connecting to the Guardian track. Some people sell "supercharged schema": hiding an instruction for AI inside the markup, inflating attributes, faking authority. Don't do it. That's hidden instruction injection, the number-one OWASP risk for AI in 2026, and it violates anti-spam policy. Worse: in models with strong safety training (like Claude), this tends to backfire. There's a 2026 study, the Injection Paradox, showing that AI detects the manipulation attempt and downranks the entire brand, quietly, with no warning. It's what you saw back in the prompt injection lesson (G.2): the hidden instruction in content is an attack, and the house trained against attacks punishes you for trying. Schema is for telling the truth in a legible way, not for tricking.
03llms.txt: know it, but don't fall in love with it
Now an item that's going to show up in every GEO article and that I need to hand you with both feet on the ground, so you don't spend energy in the wrong place. There's a proposed file called llms.txt, at the root of the site, in the spirit of robots.txt, except the idea is to summarize and point AI to what's worth reading on your site. Sounds lovely: "a map of my content made for AI".
The problem is this, and I'd rather tell you the truth than sell you hope: real adoption of llms.txt in 2026 is around 10%, and no major AI has confirmed using this file as a signal in production answers. Google publicly said it doesn't use it; and the few that touch the file (Anthropic publishes its own, Perplexity says it reads it) don't treat it as a guaranteed ranking signal. In other words, it's not standard, it's not guaranteed, and it doesn't replace anything that came before. It's "hype to watch", not a silver bullet.
For the board. llms.txt is cheap to create and doesn't hurt to have (a few minutes of work). But treat it as a low-cost experiment, not as your GEO strategy. If someone sells you llms.txt as "the secret to dominating ChatGPT", be skeptical: what actually moves the needle is the basics done well, unblocked robots, clean schema, a complete product sheet, and your brand cited by trusted third-party sources. llms.txt is the bonus, not the main course.
Notice the pattern that repeats in the AI era: the real gain is in the boring, consistent work, not the week's new trick. Every time an "magic file" or "hack nobody knows about" shows up, remember this lesson. Fair enough?
04The Golden Record: the product sheet the machine recommends
We've arrived at the most lucrative piece for whoever sells a product. Think about this: when your customer's AI agent goes to recommend (or even buy) a product in your category, it doesn't look at the pretty photo or the creative copy. It reads the SHEET: name, description, price, availability, brand, spec, all in a structured field. If your sheet is complete and clean, you're a candidate. If it's half-empty, with blank fields and vague descriptions, AI moves on to the competitor whose sheet it can read in full.
That polished sheet has a name: Golden Record. It's the product sheet with nearly all attributes filled in and structured, the so-called "agent-legible product". And the effect is impressive: stores with a well-built Golden Record get to see far more AI visibility than sparse catalogs with missing fields. It's not magic, it's AI naturally preferring the information it trusts.
The work here isn't technical, it's about organization. For every product that matters, you make sure it has:
- Name clear and official (not the internal nickname).
- Description that answers "what it is, for whom, what the benefit is", without fluff.
- Price and availability correct and up to date.
- Attributes the customer filters by (size, color, material, compatibility, whatever fits your category).
- Brand linked to your Organization entity, so AI connects the product to who you are.
And this ties directly into this track's next step, agentic commerce: protocols like ACP (OpenAI and Stripe) and UCP (Google) are defining how the agent discovers, buys, and even handles post-sale on the customer's behalf. The fuel for all of this is the product sheet. Without a Golden Record, you don't participate in the purchase AI makes, simple as that. The pretty storefront still matters for the human who clicks; the complete sheet is what matters for the machine that recommends.
05The map: why this lesson is the basement of everything
Put the four pieces together and you have your business speaking the machine's language:
- Unblocked robots in robots.txt, so AI can read (the prerequisite).
- Schema (Organization and FAQ in JSON-LD), with copyable 40-to-80-word answers (the sign it reads and cites).
- llms.txt known and maybe published, but with no illusions (the bonus).
- Golden Record of the product, complete and structured (what gets AI to recommend and, soon, buy).
Notice that none of this is creative, and all of it is decisive. It's the marketing equivalent of the "connect AI to your numbers" that finance does, or the "connect sales context" that the commercial team does: before any brilliant delivery, you connect the machine to the right source. In marketing, the source is your own site, and the consumer is the AI deciding whether to recommend you or the neighbor.
To take with you: to be found and recommended by AI, first make your business legible to the machine. Unblock the right robots in robots.txt (without this, you're not even in the game), mark up Organization and FAQ in JSON-LD with self-contained 40-to-80-word answers, treat llms.txt as a low-cost experiment (adoption ~10%, no silver bullet), and build your product's Golden Record, the complete sheet that yields much more visibility. It's the least glamorous, most lucrative work in GEO, because almost nobody does it, and never try to trick AI with hidden schema, it detects it and downranks you. Fair enough? Next up.
Now build
Your workbench for this lesson is building the Legibility Sheet for your business for AI: one page, about twenty minutes, that becomes your GEO technical to-do list. Take your real task or your own site and fill out the four fronts.
FRONT 1 · THE ROBOTS (the door)
- Open your site and go to
yourdomain.com/robots.txt(just type that into the browser). - Look for the names: GPTBot, PerplexityBot, ClaudeBot, Google-Extended. For each one, note: is it unblocked (Allow or absent) or blocked (Disallow)?
- If any is blocked and you WANT to be recommended, mark it as pending item number one for your dev or platform: "unblock this robot".
FRONT 2 · THE SCHEMA (the sign)
- Does your site have Organization schema (the brand's sheet)? And FAQ? If you don't know, note it as a question for the team.
- List the 5 questions your customer asks MOST before buying. For each one, write the answer in a 40-to-80-word block: starts by answering directly, closes the meaning on its own. These answers are the content of your FAQ schema.
FRONT 3 · llms.txt (THE BONUS)
- Decide: is it worth spending 15 minutes publishing a simple llms.txt as an experiment? Yes or no, and why. (Correct answer: you can do it, but don't count on it. It's not a priority over fronts 1, 2, and 4.)
FRONT 4 · THE GOLDEN RECORD (the sheet)
- Take your most important product. List the fields: name, description, price, availability, attributes, brand. For each, mark: is it filled in and clear, or empty/vague?
- Count how many are incomplete. That number is your AI visibility gap. Every field you close raises your chance of being read and recommended.
At the end, ask the test question: if an AI agent were about to recommend someone in my category right now, could it read my site, find my brand as a clear entity, copy an answer of mine, and describe my product without inventing anything? If the answer has a "no" on any front, you just found money on the floor. It's cheap to fix and almost nobody fixes it.
Where the trust signal that gets AI to recommend you comes from
Marking up schema and unblocking robots gets AI to READ you. But being effectively RECOMMENDED depends on one more layer, which GEO research calls the consensus signal: AI only recommends a brand with confidence when it sees it repeated, consistently, across several independent, trusted sources, not just the brand's own site. That gives us two facts that change your strategy. First, an unlinked mention in a trusted source (the so-called unlinked brand mention) correlates far more strongly with being cited by AI than the old backlink: an Ahrefs study of 75,000 brands points to around 0.66 versus 0.22. In other words, being talked about by authoritative third parties weighs more than receiving a link. Second, this explains why this lesson's technical basement is necessary but not sufficient: schema guarantees that, when AI does talk about you, it gets the facts right and cites you verbatim; consensus (PR, partnerships, presence in lists and forums AI reads, like Reddit) is what makes it WANT to talk about you. Schema is the legible sign; consensus is the reputation. You need both, and the sign is the one you control 100% today, so start there.
Practice
1. A company, wanting to 'protect content from AI', blocked GPTBot and ClaudeBot in robots.txt. What's the effect on the goal of being recommended by AI?
2. Why is writing your FAQ schema answers in self-contained 40-to-80-word blocks better than a long, tangled paragraph?
3. A vendor offers you a 'premium GEO package' centered on publishing an llms.txt and embedding hidden instructions in schema to get AI to 'prioritize' your brand. How do you assess this?
Fair enough? Let's close this lesson's point together. Marketing in the AI era has a technical basement that decides the game before any creative work shows up: you unblock the right robots in robots.txt (GPTBot, PerplexityBot, ClaudeBot, Google-Extended), because blocking them means removing yourself from the answer without noticing; you mark up your brand and your questions in schema, with 40-to-80-word answers AI copies verbatim; you know llms.txt but don't get fooled by it; and you build your product's Golden Record, the complete sheet that yields 3 to 4x more visibility and opens the door to agentic commerce. It's the least glamorous work in the track and the most lucrative, precisely because almost nobody goes down to the basement. And never, under any circumstances, try to fool the machine with a hidden instruction: the AI that recommends is the same one that detects manipulation and downranks you silently. Tell the truth, tell it in a legible way, and let the machine do the rest. Whoever tidies up the storefront for the machine today will be found tomorrow, while the competitor with the better brand stays invisible behind a locked door.
Thanks for the feedback. It helps sharpen the next lesson.