Extract the text

The words in an image, as a plain answer. extract_text fits the image to an acuity (long edge 1568 by default, never upscaled), runs tesseract on it, and answers with the text, its counts, and the engine’s name and version. No model reads the image and nothing is stored: the answer keeps the image’s sha256, width, height and byte count. The default below is the synthetic invoice from our own test corpus.


A URL is fetched by the API from a public host (no redirects, 4 MB at most); a file you pick travels as base64 and is never stored. An agent calls GET https://api.s2ar.dev/v1/extract_text?image_url=…, or the extract_text tool over MCP, and reads the text; the worked version is Extract the text. When the text has to become evidence, certify with the OCR witness seals a record.