TL;DR A metalworking manufacturer in central Italy receives customer orders as PDF email attachments. An n8n workflow discards non-order emails without calling the AI, estimates the cost of each attachment before analysing it, and has the model fill the Excel template the management software imports. The Excel lands in Google Drive; an operator imports it into the management software and confirms on Telegram. Over six months: 723 emails received, 504 discarded without calling the model, 219 orders, 300 attachments analysed at about $0.03 each.

PDF orders from email to management software: the AI fills a template, rules do the rest

Doinel Atanasiu
Doinel Atanasiu10 min read

A metalworking manufacturer in central Italy receives its customers' orders as PDF attachments to emails. I built the n8n workflow that turns them into Excel files ready for the company's management software: it filters the emails, estimates what each attachment will cost to analyse, has an AI model extract the order into a template the company defined, files everything in Google Drive and alerts the administrator on Telegram. It has been in production since 1 April 2026.

This page explains how it's built and why, with numbers from the first six months. The model does only the part no rule can do: read a PDF written by someone else and fill in the format the company decided on. The rest is rules, and the administrator configures the rules without going through me. What changed for the business (time, cost, how long it took to build) is in the article on what it really costs to automate PDF order entry instead.

The problem

Every customer sends orders in their own way, with different layouts and different units of measure. The company's management software, Danea Easyfatt, imports one specific format.

Before, the administrator downloaded each PDF to a PC, uploaded it to ChatGPT, had it generate an Excel file and checked the result, because something was often wrong. It took about 3-4 minutes per attachment, by the administrator's own estimate.

The AI was already in the process. What was missing was the process around it: a fixed output format, a rule for units of measure, an archive, and a record of what had arrived and what had been done.

One constraint shaped the choice of tools: the company already holds software licences, Google Workspace among them, and the solution had to fit those. That is why email, archive and configuration live in Gmail and Google Drive, the tools the company already uses.

Every filter before the AI costs nothing

Over six months 723 emails reached the monitored address. 219 were real orders. Had the first step been to ask the model whether an email is an order, I would have paid for an AI call on each of the other 504.

Three deterministic checks run before the model:

  1. Sender blacklist. The addresses to ignore sit in a list the administrator edits alone. Discarded: 22.
  2. Attachments and activation keyword. An email is an order if it has attachments and its subject contains one of the activation keywords, which the administration sets. Discarded: 482.
  3. Duplicates. Before analysing an attachment I check whether a file with the same name already exists in that customer's folder for the current year. Discarded: 0.

None of these checks calls a model: 70% of emails (504 of 723) stop here, at no cost.

The subject keyword has a downside I accepted: a genuine order whose subject lacks a keyword is treated as a non-order. The error stays visible, because every discarded email gets a Gmail label (more on that below).

The spending cap is decided before spending

A spending cap per attachment is a dollar threshold, set by the administrator, that the estimated cost of an analysis cannot exceed: if the estimate goes over, the document never reaches the model and a person decides whether to authorise it.

An AI extraction is billed by the token, and tokens depend on the size of the PDF. A one-page order and a PDF of many pages do not cost the same.

Before each analysis the system counts the attachment's input tokens with the provider's token-counting endpoint and adds the maximum number of output tokens the call can produce, which is the limit set on the response. It converts the total to dollars using the model's prices and compares it with the maximum spend per attachment that the administrator sets in Drive. If the estimate exceeds the threshold, that attachment is not analysed.

Counting the maximum output makes the estimate conservative: in the worst case the analysis costs what was estimated.

The threshold applies per attachment, not per email. If an email has several PDFs and only some exceed the threshold, the others are processed and the email gets the "partly too expensive" label. Over six months the cap stopped 13 emails out of 219, in whole or in part.

The cap can be overridden explicitly. For each stopped email the administrator authorises that specific email (not the sender) and forwards it again, and the system reprocesses it without applying the cap. If the email was only partly expensive, reprocessing handles the attachments skipped the first time. All 13 were authorised and reprocessed. The cap guarantees that no out-of-scale spend happens without a person having seen it.

I consider this necessary in any tool that calls a model on inputs I don't control: a very long PDF costs more and is also the most likely to produce anomalous results, and it is better to know before paying.

Process first, then the AI

Before automating anything I standardised the process with the company. The old flow left the choice of format to ChatGPT: for each PDF it decided which columns to produce, in what order and in which units.

Two things came out of that:

  • an order template: the Excel format Easyfatt imports, with fixed columns and types;
  • a unit-of-measure map: for every unit that appears in customers' orders, how the company handles it internally.

The model receives both on every extraction. It fills in the company's format and converts units with a table it didn't write.

According to the administrator, in six months no order has needed manual correction. That is their statement, not a system measurement (more on that at the end). I attribute it to how we worked: standardise the flows first, then automate them with AI.

The model is GPT-5.4 from OpenAI. I compared models from several providers on the same orders; at the time of testing it had the best balance of extraction quality and price.

Status lives where the operator already looks

An automation that doesn't show itself leaves one question hanging: "did that order arrive?". Instead of building a dashboard I put each email's status in two places the company already uses.

In Gmail, as labels. There are generic labels (error, blacklist) and one label per customer, with sub-labels that follow the order's lifecycle. The label names are in Italian, the language the company works in; English glosses are in brackets:

  • 1 - In attesa (pending)
  • 2 - Processato (processed)
  • 3 - Inserito (entered)
  • 4 - Non ordine (not an order)
  • 5 - No allegati ordine (no order attachments)
  • 6 - Parz. costoso (partly too expensive)
  • 7 - Costoso (too expensive)

The numbers in front of the names keep the sub-labels in the order of an order's journey in Gmail's sidebar. The operator sees in real time where each email stands and why one stopped.

If something goes wrong, the email gets the "errore" (error) label and both the administrator and I are notified. In six months I have never seen it fire.

In Google Drive, as archive and record. Each extracted order is saved as an Excel file next to the PDF that produced it, in a customer / year / month structure. The customer folder name is configurable. Alongside sits a tracking Excel file that records every email received: whether it was processed and, if not, why (blacklist, non-order, too expensive, error), the cost of each attachment and any errors.

The last step stays human, because the management software has no API

Danea Easyfatt cannot be driven through an API. The constraint comes from the management software, and it decided the shape of the last step: a person does the import.

To shorten that step I built a Telegram bot. When an email has been processed it gets the "processed" label, and right after that n8n sends the administrator a message with all the orders extracted from that email. Each row has two buttons:

  • Open, which opens the Excel file in Google Drive;
  • Download, which downloads the file to the PC and then changes its label to show it was downloaded.

At the bottom of the message a third button confirms that the orders from that email have been entered into the management software. It also changes its label after the click.

Two separate n8n flows sit behind the buttons. One delivers the file when the operator presses Download. The other, on confirmation, moves the email's Gmail label from "processed" to "entered". One click therefore updates both the Telegram message and the mailbox, without a second tool to keep in sync.

Whoever knows the customers shouldn't have to call a developer

Everything that changes with the life of the company sits in configuration files in Google Drive, which the administrator edits alone:

  • the sender blacklist;
  • the subject activation keywords;
  • the maximum spend per attachment;
  • the unit-of-measure map;
  • the customer folder names.

When a customer shows up with a unit of measure nobody has seen, a row goes into the map. When a sender should no longer be processed, it goes on the blacklist. None of these changes goes through me or requires touching the workflow.

The model is a node

In the workflow the model lives in a single node. Changing it takes three contained steps:

  1. replace the AI node with the new provider's;
  2. update the constants with the model name and the input and output token prices;
  3. point the cost estimate at the new provider's token-counting endpoint.

The filters, the spending cap, the template, the archive and the notifications don't change. The template and the unit map don't depend on GPT-5.4, so they stay as they are.

What I accepted losing

Keyword false negatives. An order whose subject contains none of the keywords is treated as a non-order. I preferred that risk, visible through the Gmail labels, to paying for an AI classification on 482 emails that were not orders.

Duplicates recognised by file name. The check holds if customers give their files different names from one order to the next. If the same order arrives under a different name, it isn't recognised; if a customer always reuses the same name for different orders, the second is mistaken for a duplicate. Six months without a flagged duplicate does not prove that neither case happened. The next step would be comparing the extracted order number instead of the file name, but that means paying for the extraction before finding out it's a duplicate.

Measuring corrections. The confirm button on Telegram records that an order was entered, not whether it was edited first. The "zero corrections" figure is therefore the administrator's word, not a system measurement. Measuring it would take a second button next to the confirm one.

The numbers, as of 28 September 2026

DataValue
In production since1 April 2026
Emails received723
Discarded: blacklist22
Discarded: not an order482
Order emails219
of which stopped by the spending cap13 (in whole or in part, all reprocessed)
Attachments analysed300
Errors / duplicates0 / 0
Average cost per attachment~$0.03
API usage in production~$10
API usage for development and tests~$10 (before going live)

70% of emails stopped at the deterministic checks and cost nothing. The remaining 30%, the orders, brought 300 attachments analysed at about 3 cents each: roughly $10 over six months. Another $10 or so had been spent before going live, on development and testing. The workflow reached production already tested, which is how I explain zero errors and zero duplicates.

The stack

n8n for the workflow. Gmail as the inbound channel, with its labels as visible status. OpenAI's GPT-5.4 for extraction, with the provider's token-counting endpoint for the cost estimate. Google Drive for archive, record and configuration. A Telegram bot for the last step towards Danea Easyfatt.

FAQ

Common questions and answers
Does the AI read every email that arrives?
No. Only free, deterministic checks run before the model: sender on the blacklist, presence of attachments, an activation keyword in the subject, a file already present in Drive. Over six months 504 of 723 emails were discarded this way, without spending anything on API calls.
How do you stop a very long PDF from blowing up the cost?
Before analysing an attachment the system counts its input tokens, adds the maximum output tokens the call allows, converts the total to dollars and compares it with the maximum spend per attachment set by the administrator. If it is higher, the attachment is not analysed and the email is labelled too expensive; the administrator can authorise it and have it reprocessed.
Do orders reach the management software automatically?
No. The company's management software, Danea Easyfatt, cannot be driven through an API, so the last step stays human: the administrator receives the orders extracted from each email on Telegram, downloads the Excel files, checks them, imports them and confirms the entry with a button in the same message.
How much does it cost to extract an order from a PDF with this system?
About $0.03 per attachment on average. From 1 April to 28 September 2026 the system analysed 300 attachments for roughly $10 of API usage, plus about $10 spent before going live on development and testing.
Can a different AI model be used?
Yes. The model lives in a single node of the n8n workflow. Switching means replacing that node, updating the constants with the model name and its input and output token prices, and pointing the cost estimate at the new provider's token-counting endpoint. The rest of the workflow does not change.
Interested in a system like this?
Reach out on LinkedIn, or send me an email - I read both.