Prompts / Structured JSON extraction

Data extractionjsonstructured-output

Structured JSON extraction

Extract structured data from messy text into a schema you define, with missing fields left null instead of guessed.

Copying runs entirely in your browser - nothing here is ever sent anywhere.

Fill in the variables

Extract data from the text below into JSON matching exactly this schema
(field names, types, and nesting):

{{schema}}

Rules:
- If a field's value isn't present in the text, set it to null - do not
  guess, infer from context you're not confident about, or leave the
  field out of the JSON entirely (the schema's shape must always be
  complete, even when a value is missing).
- If the text gives a value in a different format than the schema expects
  (e.g. "March 3rd" for a field typed as an ISO date), convert it, but
  only when the conversion is unambiguous - flag it as null with a note
  if it's genuinely ambiguous.
- Output ONLY the JSON, no explanation before or after it, unless a field
  explicitly asks for one.
- If the text contains multiple entities matching the schema (e.g. several
  people, several line items), return an array of objects instead of a
  single object.

Text:
{{text}}

When to use

Pulling structured data out of an email, a resume, a support ticket, a scanned-and-OCR’d document, or any other free-text source you need to turn into a database row or API payload. Also useful as a first pass before building a proper extraction pipeline, to see what the schema needs to handle.

Why it works

The most damaging extraction failure isn’t a missing value - it’s a confidently invented one that looks correct and gets written to a database as if it were real data. Explicitly instructing null-over-guessing, and “only convert format when unambiguous,” trades a small amount of recall for a large amount of precision, which is almost always the right tradeoff for structured data feeding a downstream system.

Variations

  • Add “This came from OCR and may have character-recognition errors - correct obvious OCR mistakes (0/O, 1/l) but flag anything you had to guess at” for scanned documents.
  • For a schema with an enum field, list the exact allowed values in the schema description and add “if none of the allowed values fit, use null” rather than letting the model invent a new one.
  • Ask for a confidence field per extracted value (high/medium/low) if a human will be spot-checking the output before it’s trusted.