What leaves your machine, and what does not
The invoice file is the thing people worry about. It does not leave the browser. This page is the rest of the picture: the few requests that do hit a server, who sees them, and how long they last.
Controller
Tallystick is run by a single operator outside the EU, contactable at info@tally-stick.com. The full legal identity is published on the German imprint, which is where a controller statement is legally required. There is no data-protection officer, because this is a one-person operation and none is required.
The invoice file
Parsing runs in your browser with JavaScript. The file is not uploaded and not copied to object storage. That holds for every tool on this site, including the reading of a scanned PDF described next. There are two things a file can leave behind, both listed in full further down: if a read fails, the error log keeps a description of the reading and the first lines the reader got off the page, and if you press Send the file on that error, the file itself goes to the repair inbox.
The one exception is a report you send yourself: when a file cannot be read, the page offers to attach the exact file to the error report so the reader can be repaired. That offer sits under the error message wherever a file is dropped, and on the format converter when a conversion fails. The attachment travels only when you press that button, goes to the operator's repair inbox, is read only to repair the reader, is never sold, never shared, and is removed once the repair is done.
Reading a scanned PDF
PDF to XRechnung fills the form from an ordinary invoice PDF. A PDF with a text layer is read straight out of the file. A PDF that is only a picture of an invoice, a scan or a photo, is passed through text recognition instead, and that recognition also runs in your browser tab.
To do it the page fetches two open-source libraries, pdf.js and Tesseract.js, together with a recognition language pack, from the jsDelivr CDN when you drop the file. That fetch tells jsDelivr your IP address and user agent, the way loading any third-party script does. Your invoice is not part of the request: what travels is the request for the library, and the reading happens afterwards on your machine. If you would rather not involve a CDN at all, the desktop app does the same job with nothing fetched at run time.
What a server does see
- Page requests. Cloudflare hosts the site. Their edge logs a request the way any HTTPS host does (IP, URL, user agent, time).
- Visit counting. Cloudflare Web Analytics runs on every page: a small script that reports the URL, the referrer, and how fast the page loaded. It sets no cookie, reads no cookie, and builds no fingerprint or identifier, so one visitor cannot be followed from page to page or from one day to the next. It exists because the host's own request log counts crawlers as visitors, and by that measure most of this site's traffic is not human. No part of an invoice is involved: the script sees the address of the page, never what you opened on it. Cloudflare is already the host, so this adds no new company to the list below. If your browser blocks it, nothing on the page breaks and the visit is simply not counted.
- Conversion counter. After a successful parse the page posts a number (how many files just succeeded, capped) to
/api/count. The server stores one integer. No file name, no invoice, no cookie. - Page and time measurement. The owner also needs to know which pages earn their keep, so when a page is opened or hidden it posts to
/api/visit: the page's address, the hostname of the page that linked here (never the full URL), and on the way out how many seconds it stayed open. Days are aggregated into totals per page, five coarse time buckets, and a rough class of what kind of site linked here. To estimate visitors per day without storing anyone, the address is hashed together with a server-side key and the day's date, and the hash is used only to set two bits in a counting filter: it is never written down, cannot be turned back into an address, and no cookie or identifier of any kind is set. If the request never arrives, the page works exactly the same. - Failure counting. When a file cannot be read, the page posts to the same endpoint which tool was open, a fixed code naming the check that refused the file, the file extension, and a size band. It is there so a format that never works can be found and fixed without waiting for anyone to report it, and a free tool is the last place a defect gets reported. The codes are a closed list: no file name and no invoice content. The parsing itself still happens only in your browser and the file is still never uploaded.
- The error log. A code says that a file failed and nothing about why, and a defect nobody can see is a defect nobody fixes. So when a read fails, and only when it fails, the page also posts to
/api/error: the message you were shown, how many pages the document has, how many words were on it, whether it had to be read by OCR, how long the read took, which program wrote the file according to its own metadata, and up to the first 4000 characters the reader managed to read off the page. On an invoice that failed to parse, that last part is invoice text, and it is the only thing that ever says which label or layout defeated the reader. It is not the file: no attachment, no image, no second page, and no file name. Rows are kept for up to a year, are read by one person to repair the reader, are never sold and never shared. Write to the address above to have one deleted. - Feedback you send. If you use the form, the server stores the message, an optional email, a short technical context (error text, file name, file size), the page URL, your user agent, and a timestamp. The invoice itself is not attached. Copies live in a key-value store for up to one year and are emailed to the address above so the operator can answer.
- Theme preference. Light or dark is saved in
localStorageon your device. It is not sent to the server. - Signed EN 16931 record. Only when a Team subscriber presses the button. The page posts a SHA-256 of the file, the ruleset id, the finding codes and counts, and the invoice number and issue date to
/api/sub/report. The invoice file itself is not sent. The server signs that digest and returns it; it stores nothing from the request.
Cookies
The converter does not set a tracking cookie. Cloudflare may set cookies that belong to the host (for example to distinguish a human from a bot). They are not used to identify anyone.
Legal basis and retention
Page hosting, the conversion counter and the page-and-time measurement are required to operate the service (legitimate interest in keeping a free tool online and honest about how much it is used). Feedback is processed because you asked for a reply or reported a defect (consent / contract-like request). You can write to the address above to have a stored feedback row deleted. Host logs follow Cloudflare's own retention.
Processors
Cloudflare, Inc. provides DNS, TLS, hosting, the visit counting described above, and the key-value store. Mail delivery of feedback goes through a small Worker that can send only to the operator address. There is no advertising network and no invoice is shared with anyone.
Your rights
If the GDPR applies to a processing described above, you can ask for access, correction, deletion, restriction, or a copy, and you can complain to a supervisory authority. Write to info@tally-stick.com.