
PDFmdx – highly automated middleware for Intelligent Document Processing (IDP), which analyses unstructured PDF data streams using layout and content classification, dynamically partitions them (splitting), extracts structured metadata via zonal OCR and relative anchor fields, and makes this metadata available as XML or CSV for forwarding to third-party systems (ERP/DMS). Through native integration with the PDF2Printer printing engine and an integrated SMTP module, this digital workflow can be seamlessly, rule-based, and fully automatically transferred to physical document output or digital delivery:
• Document classification & splitting: automatic detection of document type based on content and dynamic splitting (segmentation) of composite PDFs based on separation criteria such as text breaks or detected barcodes.
• Multi-format Barcode recognition: automatic decoding of 1D and 2D barcodes (e.g., Code 128, QR Code, DataMatrix) to control splitting processes, for file naming, or as a valuable index value for metadata export.
• Intelligent data extraction (Parser Engine): extraction of specific metadata (e.g., invoice amounts or customer numbers) using geometrically flexible anchor fields and regular expressions (RegEx).
• Dynamic finishing: content-driven application of digital letterhead (overlay/underlay), electronic signatures, or automatic removal of blank pages (Blank Page Detection).
• Automated email sending (SMTP output): rule-based, direct sending of processed documents as email attachments, with the recipient’s address, subject line, and message text dynamically generated from the previously extracted metadata (e.g., email address from the invoice header).
• Event-Driven file monitoring & print control: the PDF2Printer windows service continuously monitors the output directories of PDFmdx and immediately triggers the spooling process, including specific parameters such as printer assignment, targeted tray selection, and number of copies.
• Lifecycle management & error handling: after successful printing or sending, the system moves the documents to an archive directory, while failed jobs are isolated and placed in dedicated error folders for administrative review.