Unstructured data
Turn PDF tables into data you can model
Invoices, statements and reports carry real numbers in layouts built for a human reader. Nexadata finds the tables inside them, stitches the ones that run across pages, and hands you rows.
At a glance
Read-onlyDocuments this handles
- Invoices, statements and remittances
- Board, fund and regulatory reports
- Manifests, schedules and inspection results
Setup
Once per document layout, not once per document.
- SOC 2 Type II
- Zero data retention
- SSO / SAML / RBAC
- Cloud-agnostic
The platform processes your data. An LLM never touches it.
The connector
What Nexadata does with a PDF
It finds the tables, including the awkward ones
Nexadata reads the document and detects every table in it, stitching back together any that continue across a page break. A line-item table running over three pages arrives as one continuous table rather than three fragments you have to reassemble.
You decide what each table is
Review what was found, rename the tables that matter, rotate the ones whose headers run down the side instead of across the top, and discard the rest. Then nominate the main table and layer on the joins and transformations that turn it into something useful.
Do it once per layout, not once per document
Every choice you make is recorded as a template. Next quarter's statement in the same layout is processed against it, with the same tables recognized, named, oriented and joined the way you decided, and you get a new dataset without repeating the review.
Length stops being the problem
A 600-page quarterly report is as legitimate an input as a two-page invoice. Curating something that size is real work the first time and close to free every quarter after, which is what turns a recurring pile of documents into a dependable input rather than a standing task.
How it works
From document to dataset
Connect, transform, map, review. Each step is guided by a no-code copilot, and you approve the plan before anything runs.
- 01
Connect
Pick the file-based connection holding the document, choose PDF as the data format, and select a representative file rather than an unusual one, because this first pass defines the template.
- 02
Transform
Stitch the tables you kept into one picture, join them to the systems the document is about, and reshape the result by describing it to the Transform Copilot.
- 03
Map
Match the supplier names, account codes and line descriptions printed in the document to the members your model already uses, including the ones written differently every time.
- 04
Review
Check the columns and the plan before anything runs, and confirm the extraction did what you expected before it becomes a scheduled input.
Use cases
What people run through it
Supplier and freight invoices
Line items locked in a monthly stack of PDFs, lifted into rows you can allocate, accrue and plan against, without anyone re-keying them into a spreadsheet first.
See the use caseFund and portfolio reports
A long quarterly report curated once, then run every quarter after as the numbers change but the layout does not.
See the use caseBank and brokerage statements
Recurring statements turned into a dependable feed for reconciliation and cash reporting, on a schedule rather than on someone's to-do list.
Read the guideSet it up
Step-by-step documentation
Every screen, in order, with screenshots.
Setting up a PDF dataset
The full Process PDF flow, from table detection to a finished dataset.
Read the guideSupported data formats
How tabular, spreadsheet and PDF formats differ and when to use each.
Read the guideHow to create a dataset
The wizard every dataset goes through, whatever its format.
Read the guideNexadata Hub
Managed storage, if your documents have nowhere else to live yet.
Read the guideQuestions
PDF data extraction FAQ
- What kinds of PDF does this work on?
- Documents with tables in them: supplier and freight invoices, bank and brokerage statements, utility bills, remittance advices, shipping manifests, lab and inspection results, insurance schedules, fund and board reports, regulatory filings. Nothing about the flow is specific to an industry.
- What happens when a table runs across several pages?
- It is stitched back into one continuous table. A line-item table spanning three pages arrives as a single table rather than three fragments for you to reassemble.
- Do I have to set this up for every document?
- No, once per layout. Your decisions about which tables matter, how they are oriented and how they are joined are saved as a template, and the next document in that same layout runs against it without repeating the review. What a template cannot absorb is a structural change, such as a vendor redesigning their invoice.
- Is there a page limit?
- Length is not the constraint people expect. A 600-page quarterly report is as legitimate an input as a two-page invoice, and because the curation is per layout rather than per document, the long ones are where this pays off most.
- Is the document processed by a third party?
- Yes, and it is worth being precise about which. Text and table extraction runs on Amazon Textract, an AWS document-analysis service. Textract is not a language model, so no LLM reads your document, which is why the line above still holds. Nexadata runs with the AWS AI services opt-out applied, which means your documents are permanently deleted immediately after processing, no copy of the output is retained, and nothing is used to improve AWS models. Everything after extraction happens inside the Nexadata platform. Amazon Textract data privacy, from AWS
More integrations
Other ways to connect
Pigment
Read from and write to Pigment blocks, views and lists, with column validation on every writeback.
ExploreAnaplan
Reach past your models into Anaplan's own tenant data, and load results back through the actions your model already runs.
ExploreHubSpot
Read any standard or custom object with its associations, and write results back with upsert, create or update.
ExploreSalesforce
Pull via SOQL or saved reports, tune extract mode for volume, and write back with insert, update, upsert or delete.
ExploreConnector Copilot
Load a spec, authenticate, describe the data you want, and get a typed dataset. Over 20,000 public APIs are documented in OpenAPI.
ExploreSpreadsheets
Point at the sheet and the cell where the real data starts, and load a clean typed table out of a messy workbook.
ExploreSee it on your data
Start free with your first use case, or talk to us about your stack.