Schema Markup Validator: How to Validate Structured Data for AI Search

A schema markup validator tells you one thing: whether machines can parse what you wrote. It does not tell you whether Google will show a rich result. Those are two different questions, answered by two different tools, and confusing them is why most structured data audits go sideways.

Here’s the short version. Use validator.schema.org for spec compliance. Use Google’s Rich Results Test for feature eligibility. Use Search Console’s Enhancements reports for what Google actually indexed on your live site. All three, in that order.

Below: what each tool catches, what each one misses, how to read errors versus warnings, the mistakes that break markup silently, how to validate thousands of URLs without clicking through a form, and the honest answer on whether any of this helps you get cited by AI.

Schema Markup Validator vs Rich Results Test vs Search Console

Three tools, three jobs. They overlap just enough to confuse people and differ just enough to matter.

ToolValidates againstCatchesMisses
Schema Markup Validator (validator.schema.org)The full Schema.org vocabularySyntax errors, unrecognised types and properties, wrong value shapes, the whole extracted data graphGoogle’s required properties, rich result eligibility, anything about your live index
Rich Results TestGoogle’s documented rich result featuresMissing required properties, feature eligibility, JavaScript-rendered markup, the HTML Googlebot actually sawSchema types Google has no feature for, content mismatches, whether the result ever displays
Search Console EnhancementsWhat Googlebot already crawled on your live siteSite-wide error counts by type, sample affected URLs, trends, fix re-validationAnything not yet crawled, staging URLs, code snippets, types with no report
Unparsable structured data reportRaw parseabilityBlocks so broken Google can’t tell which type you meantEverything else. No warnings, no valid items, errors only

What the Schema Markup Validator does

It’s the old Structured Data Testing Tool with the Google-specific bits removed. Google retired the SDTT, handed generic validation to the Schema.org community, and still hosts the tool. It extracts JSON-LD 1.0, RDFa 1.1 and Microdata, draws the data graph it found, and flags syntax mistakes.

That’s the strength and the trap. It checks against all of Schema.org, so it will cheerfully approve markup for a type Google has never supported. A green result in the Schema Markup Validator means “this is valid Schema.org.” It does not mean “Google will do anything with this.”

What the Rich Results Test does

It answers one question: which Google rich result features can this page generate? It renders the page the way Googlebot does, JavaScript included, and shows you the rendered HTML it used. If a required property is missing, it says so and the feature drops off the list.

Its blind spot is scope. Paste in a page with perfectly good Organization or Dataset markup and you may get “No items detected” for the feature you were hunting. Eligibility isn’t display either. Google decides per query whether to show the enhanced result at all.

What Search Console tells you that neither test can

Both testers look at one URL, right now. Search Console looks at your whole site, as Google last crawled it. That gap is everything once you’ve shipped. Enhancements reports split errors and warnings by type, sample the affected URLs, and let you request re-validation after a fix.

The catch is lag. Reports reflect crawl, not deploy. Expect days on a busy site and weeks on a slow one. The Unparsable structured data report sits outside the per-type reports for a reason: those blocks failed before Google could classify them at all.

How to read a schema markup validator report: errors vs warnings

Errors block. Warnings don’t. That’s the rule, with three asterisks.

An error in the Rich Results Test means a required property is missing or holds the wrong value type. The page is not eligible for that feature, full stop. A warning means a recommended property is absent. You stay eligible, you just hand Google less to work with.

The Schema Markup Validator speaks a different language entirely. It reports syntax failures and properties that don’t belong on the type you declared. It has no concept of “required” at all, because required is a Google policy layer, not a Schema.org one. A clean pass in one tool tells you nothing about the other.

The three asterisks

  • Some warnings behave like errors. A merchant listing missing priceCurrency stays technically eligible and still fails to show a price. Anything that feeds the visible snippet is effectively required.
  • “Valid with warnings” is fine. Search Console still counts those items. Don’t burn a sprint clearing yellow triangles on pages that make no money.
  • Unparsable beats both. If a block lands in the Unparsable report, none of its properties exist as far as Google is concerned. Fix those first, every time.

Triage in that order: unparsable blocks, then errors grouped by template, then warnings ranked by revenue. Grouping by template is the step people skip. One broken partial in a product layout shows up as 40,000 affected URLs in Search Console and one fix in one file.

One more framing that saves arguments: neither a schema markup validator pass nor a Rich Results Test pass is proof of anything appearing in search. They’re gates, not guarantees.

Seven mistakes that quietly break structured data

Most broken markup validates fine. These are the failure modes that survive a green checkmark.

  1. Markup that doesn’t match visible content. The biggest policy risk, and the one thing no validator will ever flag. Both tools will happily approve a 4.8 aggregate rating that appears nowhere on the page. Google treats this as spam, and it’s manual-action territory. Every marked-up value needs a visible counterpart.
  2. The wrong @type. Article on a category page. LocalBusiness on a page that isn’t a location. Product on a grid of forty products. The syntax passes, Google shrugs, nothing happens and nothing tells you why.
  3. Missing required properties. Product without offers, review or aggregateRating. Event without startDate or location. JobPosting without datePosted. Only the Rich Results Test or a crawler carrying Google’s ruleset will catch these.
  4. Broken @id references. @id is a pointer, not a label. When the node it points to isn’t on the page, or the string drifts (trailing slash, http versus https, www versus bare), you get orphan nodes and a fragmented graph. Pick a canonical pattern like https://example.com/#organization and never touch it again.
  5. Multiple conflicting blocks. Theme, SEO plugin and page builder each emit an Organization node with a different logo and a slightly different name. Google picks one, or picks none.
  6. JSON-LD syntax errors. A trailing comma. Smart quotes pasted out of a word processor. An unescaped double quote inside a description. One bad character kills the entire block, not just that property.
  7. Client-side injection that never renders. Schema written by a tag manager exists for the Rich Results Test, which renders JavaScript. It may not exist for a curl-based CI check, and in production it depends entirely on render budget.

Notice that only three of those seven are things a tool can flag automatically. The rest need a human comparing markup to page.

A step-by-step schema markup validator workflow

Ten steps, in order. Follow them and you’ll catch nearly everything before it reaches production.

  1. Write the markup as JSON-LD in a single <script type='application/ld+json'> block. Google prefers it, and it’s the only format you can lint independently of the page.
  2. Run the raw JSON through any JSON linter first. Commas and quotes die here, in two seconds, before an SEO tool is involved.
  3. Paste the snippet into the Schema Markup Validator using code mode rather than URL mode. Clear every syntax error and every unrecognised property.
  4. Check your @type against Google’s search gallery. If the type isn’t listed there, no rich result exists for it. Decide consciously whether you still want the markup.
  5. Run the Rich Results Test against the staging URL. Confirm the specific feature you want is detected, not just that “something” was found.
  6. Read every marked-up value against the rendered page. Price, rating, author, date, availability. If a value isn’t visible, either delete it or publish it.
  7. Ship to one URL first. One template, one page, not the whole site.
  8. Open URL Inspection in Search Console, run the live test, and confirm the rendered HTML contains your block.
  9. Wait for a crawl. Three to fourteen days is normal. Then check the Enhancements report for that type, plus the Unparsable report.
  10. Roll out across the template, re-crawl the full site with a desktop crawler, and diff the item counts against what you expected.

Two habits make this stick. Re-run a schema markup validator check after any theme, plugin or framework update, because that’s exactly when duplicate blocks appear. And keep a saved list of your ten highest-value URL patterns so a spot check takes five minutes instead of an afternoon.

Validating structured data at scale

Pasting URLs into a form caps out around twenty pages. Past that you need a crawler, a pipeline, or both.

Crawl-based validation

Screaming Frog SEO Spider validates JSON-LD, Microdata and RDFa in one pass, against both the Schema.org vocabulary and Google’s rich result requirements, and it separates errors from warnings the same way the official tools do. In practice a desktop crawler is just a bulk schema markup validator with a spreadsheet attached. Turn on JavaScript rendering if any of your markup is injected client-side, expect crawls to run several times slower, and budget for a paid licence.

Sample by template, not by URL. Five product pages, five category pages, five blog posts and five location pages will surface almost every systemic problem on a 500,000-URL site. Random sampling mostly rediscovers the same product template forty times.

Put a schema markup validator step in CI

There’s no public Google validation API, and scraping the Rich Results Test is a bad foundation to build on. Build something smaller and more reliable instead:

  • At build time, extract every application/ld+json block and run JSON.parse on it. Fail the build on any parse error. This one check kills the most common production incident.
  • Write per-template assertions as ordinary unit tests. Product templates must have offers.price, offers.priceCurrency and offers.availability. Article templates must have headline, datePublished and author.name.
  • Snapshot your @id graph and diff it on every deploy. A domain migration or a trailing-slash change will silently orphan half your nodes otherwise.
  • Warn, don’t fail, on missing recommended properties. Otherwise the team learns to skip the check entirely.

For ongoing monitoring, the Search Console API returns rich result performance by type, so you can alert on a sudden drop instead of discovering it a month later. Pair that with a monthly crawl and you have coverage without anyone opening a testing form.

Which schema types still matter in 2026

Google’s supported list moves. Two major types left it entirely, and one disappeared quietly in 2024.

Schema typeGoogle rich result statusWorth implementing in 2026?
Article, NewsArticle, BlogPostingSupportedYes, headline, image and date treatments
BreadcrumbListSupportedYes, cheap and sitewide, changes the URL line
Product with OfferSupportedYes, price, availability, merchant listings
Review snippet, AggregateRatingSupported on allowed host types onlyYes, if the ratings are real and visible
OrganizationNo visual snippetYes, knowledge panel, logo, entity disambiguation
LocalBusinessSupportedYes for any physical location
EventSupportedYes, one of the strongest remaining features
JobPostingSupportedYes for employers and job boards
RecipeSupported, including carouselsYes
VideoObjectSupported, key moments and clipsYes if video is the point of the page
QAPageSupportedOnly for genuine user-answered pages
DiscussionForumPostingSupportedYes for forums and user-generated content
ProfilePageSupportedYes for author and creator pages
DatasetSupported via Dataset SearchNiche but valuable
SoftwareApplicationSupportedYes for app pages
CourseSupported, course list and course infoYes for education
ImageObject with IPTC metadataSupported, licensable badgeYes for stock and photo libraries
SpeakableLimited beta, news publishers onlyLow priority
FAQPageDeprecated 7 May 2026No rich result. Keep only if it structures real content
HowToDeprecated September 2023No
WebSite sitelinks search boxRemoved 21 November 2024No

The FAQ deprecation, precisely

Google’s own documentation is blunt: “As of May 7, 2026, FAQ rich results are no longer appearing in Google Search.” The rest of the timeline matters if you have tooling wired up. The FAQ search appearance and rich result report were dropped in June 2026, along with support in the Rich Results Test, and Search Console API support for the FAQ rich result ends in August 2026.

HowTo went first. Google removed the documentation on 14 September 2023, after the rich result stopped showing on both desktop and mobile. The sitelinks search box followed on 21 November 2024. The pattern is consistent: Google keeps the features that change the SERP layout in ways users engage with, and drops the ones that mostly got gamed.

None of this makes FAQPage invalid. It’s still real Schema.org, a schema markup validator will still pass it, and it still describes your content accurately for every surface that isn’t Google Search. Just don’t build a content plan around a SERP accordion that no longer exists. If you’re generating question blocks, generate them because the questions are good. Our FAQ schema generator outputs valid markup either way.

Does schema markup help AI citations? The honest answer

Directly? The evidence says barely, or not at all.

Start with Google. Its documentation on AI features states plainly: “There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” On markup specifically, Google Search Central says “There’s also no special schema.org structured data that you need to add.” That is about as unambiguous as Google gets on anything.

Then the data. In May 2026, Ahrefs published a difference-in-differences study of 1,885 pages that added JSON-LD between August 2025 and March 2026, matched against roughly 4,000 control pages. The measured effect on citations was −4.6% in AI Overviews, +2.4% in AI Mode and +2.2% in ChatGPT. The last two were statistically indistinguishable from zero. Their conclusion: “Adding schema produced no major uplift in citations on any platform.”

One caveat is worth keeping. Those pages already had 100+ AI Overview citations before the change, so the study says little about a page starting from zero visibility. And no schema markup validator can tell you whether an LLM will cite you. That’s not what it measures.

Where structured data still earns its place

Entity resolution. Organization markup with a stable @id, a logo, a description and accurate sameAs links gives Google’s Knowledge Graph a clean node for your brand. Google says the markup helps it “disambiguate your organization in search results,” and the Knowledge Graph is part of what grounds AI Overviews and AI Mode. Indirect, but real. There’s no required property list to satisfy either, since Google states Organization has no required properties. You add what applies, on your homepage or about page only.

The second reason is cheaper still. LLM crawlers fetch raw HTML. A JSON-LD block sits in that HTML as an unambiguous statement of who published what, when, and about whom, without the model having to infer it from prose. It costs nothing to be legible.

What structured data for AI search will not do is substitute for being the clearest source on the page. If you want to know where you actually stand, run a GEO audit, read the fundamentals in our LLM SEO guide, and check whether ChatGPT cites your site before you spend another week on markup.

WordPress schema markup: the duplicate block problem

Most WordPress schema markup problems aren’t missing markup. They’re too much of it.

Yoast SEO and Rank Math both output a full @graph: Organization, WebSite, WebPage, Article, Person, all cross-referenced by @id. Your theme may add its own. A page builder may add a third set. A review plugin adds a fourth. Nothing errors. You just end up with three Organization nodes carrying three different logos and Google choosing between them.

The check takes a minute. View source, search for ld+json, count the blocks. Then run the URL through a schema markup validator and count the top-level nodes it draws. Two Organization entities or two WebSite entities means you have a conflict to resolve.

The fix is boring. Pick one source of truth and switch schema output off everywhere else. Yoast and Rank Math both expose settings for this, most modern themes have a toggle, and hand-written blocks should reuse the plugin’s existing @id values rather than inventing new ones. Different types of schema markup can coexist happily on one page as long as they reference each other instead of competing. If you’re building custom markup from scratch, our schema generator produces a single clean block you can drop in and validate.

Frequently asked questions

Is the Schema Markup Validator the same as the old Structured Data Testing Tool?

Effectively yes. Google retired the Structured Data Testing Tool and moved generic validation to validator.schema.org, which Google still hosts as a service for the Schema.org community. The difference is that all Google-specific checks were stripped out, so it now validates against the Schema.org vocabulary only.

Why does the Rich Results Test find nothing when validator.schema.org shows my markup fine?

Because they answer different questions. A schema markup validator confirms your markup is valid Schema.org; the Rich Results Test only reports types Google has a rich result feature for. If your type isn’t in Google’s search gallery, “no items detected” is the expected answer, not a bug.

Do I need to fix warnings, or only errors?

Fix errors first, because they make the page ineligible for the feature. Warnings mean a recommended property is missing and the page stays eligible. The exception is any property that feeds the visible snippet, like priceCurrency on a merchant listing. Treat those as required.

Is FAQ schema still worth adding in 2026?

Not for Google rich results. FAQ rich results stopped appearing on 7 May 2026, the report and Rich Results Test support were dropped in June 2026, and Search Console API support ends in August 2026. FAQPage is still valid markup and still describes your page honestly, so keep it where it reflects real questions and answers.

Does schema markup get you cited by ChatGPT or AI Overviews?

There’s no evidence it does directly. Google states there’s no special schema.org structured data needed for AI Overviews or AI Mode, and Ahrefs’ May 2026 study of 1,885 pages found no meaningful citation uplift on any platform. The real value is entity clarity for the Knowledge Graph, not citation volume.

How long before Search Console shows my new structured data?

Usually three to fourteen days, because the reports reflect Google’s last crawl rather than your deploy. Use URL Inspection’s live test for immediate confirmation on a single page. If nothing appears after a month on an indexed URL, check the Unparsable structured data report.

Can a page have more than one JSON-LD block?

Yes, Google combines multiple blocks on the same page. The problem isn’t the count, it’s contradiction. Two Organization nodes with different names or logos force Google to guess. Use @id to link nodes together rather than repeating the same entity.

Does Google read schema markup added by JavaScript?

Yes. Google renders pages before extracting structured data, and the Rich Results Test shows you the rendered HTML it used. It’s still slower and more fragile than server-rendered JSON-LD, and it’s invisible to CI checks that only fetch raw HTML. Server-render it where you can.

What’s the difference between @id and url in schema markup?

url is a property describing where something lives. @id is the unique identifier for a node in the graph, used so other nodes can point at it. They’re often the same string, but @id must be stable and unique across your site. Change it and you break every reference aimed at it.

Do I need schema markup on every page?

No. Organization markup belongs on your homepage or a single about page, per Google’s guidance, not sitewide. Page-level types like Article, Product or Event go on the pages that actually describe them, and breadcrumbs go everywhere. Blanket-applying one type is a common cause of wrong-@type problems.

What’s the fastest way to validate structured data across a whole site?

Crawl it. Screaming Frog SEO Spider validates against both the Schema.org vocabulary and Google’s rich result requirements in one pass and separates errors from warnings. For prevention rather than detection, add a build-time JSON.parse check plus per-template property assertions in your CI pipeline.

zulqarnain, founder of LLM Optimization

Written by

zulqarnain

Writes about how AI search engines such as ChatGPT, Google AI Overviews, Perplexity, Gemini and Claude choose the sources they cite.

Scroll to Top