All integrations

Unstructured data

Turn PDF tables into data you can model

Invoices, statements and reports carry real numbers in layouts built for a human reader. Nexadata finds the tables inside them, stitches the ones that run across pages, and hands you rows.

At a glance

Read-only

Documents this handles

  • Invoices, statements and remittances
  • Board, fund and regulatory reports
  • Manifests, schedules and inspection results

Setup

Once per document layout, not once per document.

  • SOC 2 Type II
  • Zero data retention
  • SSO / SAML / RBAC
  • Cloud-agnostic

The platform processes your data. An LLM never touches it.

The connector

What Nexadata does with a PDF

It finds the tables, including the awkward ones

Nexadata reads the document and detects every table in it, stitching back together any that continue across a page break. A line-item table running over three pages arrives as one continuous table rather than three fragments you have to reassemble.

You decide what each table is

Review what was found, rename the tables that matter, rotate the ones whose headers run down the side instead of across the top, and discard the rest. Then nominate the main table and layer on the joins and transformations that turn it into something useful.

Do it once per layout, not once per document

Every choice you make is recorded as a template. Next quarter's statement in the same layout is processed against it, with the same tables recognized, named, oriented and joined the way you decided, and you get a new dataset without repeating the review.

Length stops being the problem

A 600-page quarterly report is as legitimate an input as a two-page invoice. Curating something that size is real work the first time and close to free every quarter after, which is what turns a recurring pile of documents into a dependable input rather than a standing task.

How it works

From document to dataset

Connect, transform, map, review. Each step is guided by a no-code copilot, and you approve the plan before anything runs.

  1. 01

    Connect

    Pick the file-based connection holding the document, choose PDF as the data format, and select a representative file rather than an unusual one, because this first pass defines the template.

  2. 02

    Transform

    Stitch the tables you kept into one picture, join them to the systems the document is about, and reshape the result by describing it to the Transform Copilot.

  3. 03

    Map

    Match the supplier names, account codes and line descriptions printed in the document to the members your model already uses, including the ones written differently every time.

  4. 04

    Review

    Check the columns and the plan before anything runs, and confirm the extraction did what you expected before it becomes a scheduled input.

Questions

PDF data extraction FAQ

What kinds of PDF does this work on?
Documents with tables in them: supplier and freight invoices, bank and brokerage statements, utility bills, remittance advices, shipping manifests, lab and inspection results, insurance schedules, fund and board reports, regulatory filings. Nothing about the flow is specific to an industry.
What happens when a table runs across several pages?
It is stitched back into one continuous table. A line-item table spanning three pages arrives as a single table rather than three fragments for you to reassemble.
Do I have to set this up for every document?
No, once per layout. Your decisions about which tables matter, how they are oriented and how they are joined are saved as a template, and the next document in that same layout runs against it without repeating the review. What a template cannot absorb is a structural change, such as a vendor redesigning their invoice.
Is there a page limit?
Length is not the constraint people expect. A 600-page quarterly report is as legitimate an input as a two-page invoice, and because the curation is per layout rather than per document, the long ones are where this pays off most.
Is the document processed by a third party?
Yes, and it is worth being precise about which. Text and table extraction runs on Amazon Textract, an AWS document-analysis service. Textract is not a language model, so no LLM reads your document, which is why the line above still holds. Nexadata runs with the AWS AI services opt-out applied, which means your documents are permanently deleted immediately after processing, no copy of the output is retained, and nothing is used to improve AWS models. Everything after extraction happens inside the Nexadata platform. Amazon Textract data privacy, from AWS

See it on your data

Start free with your first use case, or talk to us about your stack.