Skip to main content

Purchase Document Extractor Port

Source-backed guide for the intake extractor contract in apps/backend/src/purchase-document-ingestion/.

Use this page when adding a new source (PDF OCR, spreadsheet, certifier pull). For the module workflow and HTTP surface, see Purchase Document Ingestion.


Intent​

Every extractor produces the same source-agnostic ExtractedPurchaseDocument. Downstream use cases (supplier match, review flags, line resolution, posting) never branch on XML vs PDF vs future formats.

That is the same pattern as FEL's ElectronicCertificationProvider and ecommerce's EcommerceProvider: domain shape in the port, HTTP/XML details in the adapter.


Hexagonal map​

LayerCodepathsResponsibility
Domaindomain/purchase-document-extractor.port.tsCanonical extracted header/lines, PurchaseDocumentExtractorPort
Applicationapplication/extract-purchase-document.use-case.tsLoad stored bytes, call adapter, persist lines/flags, match supplier
Infrastructureinfrastructure/extractors/fel-xml-extractor.adapter.ts, purchase-document-extractor.registry.tsFEL DTE XML parse, MIME routing
Interfacesnone for extractionExtraction is never on the request path

Port contract​

export interface PurchaseDocumentExtractorPort {
supports(mimeType: string, filename: string): boolean;
extract(input: {
buffer: Buffer;
filename: string;
businessTaxId: string;
}): ExtractedPurchaseDocument | Promise<ExtractedPurchaseDocument>;
}

businessTaxId is the ingesting business NIT (from business.taxId), not the supplier. It exists so adapters can enforce the receiver-NIT check (FR-007).

ExtractedPurchaseDocument.source is currently "fel_xml" | "pdf" | "spreadsheet" | "fel_pull". Only fel_xml is implemented. PDF-only files never call an extractor; the use case sets unparsed instead.


Runtime adapter​

FelXmlExtractorAdapter implements the port today.

ConcernAdapter behavior
MIMEapplication/xml, text/xml, or filename ending .xml
Parserfast-xml-parser with removeNSPrefix: true so SAT:/DTE:/sat: prefixes do not matter
TimezoneGuatemala fixed UTC-6 (-06:00); no TZ database
DTE typesFACT, FCAM, NCRE, NDEB only; others throw UnsupportedDocumentTypeError
Receiver NITmismatch throws DocumentRecipientMismatchError — the use case deletes the draft
Arithmeticquantity × unit price vs line total collected as validationIssues (flags, not hard fail)

PurchaseDocumentExtractorRegistry exists to pick an adapter via supports(). Extraction currently calls FelXmlExtractorAdapter directly after confirming an XML documentLink exists. When a second adapter ships, route through the registry instead of adding source branches in the use case.


Adding an extractor​

  1. Implement PurchaseDocumentExtractorPort under infrastructure/extractors/.
  2. Map provider-specific fields into ExtractedPurchaseDocument only — do not persist certifier XML paths on the ingestion row from the adapter.
  3. Register the adapter in PurchaseDocumentExtractorRegistry.
  4. Keep receiver-NIT, unsupported-type, and arithmetic rules as domain errors / validationIssues, not HTTP status codes.
  5. Do not run extraction in the controller. Enqueue purchase-document-extraction / extract like the XML path.
  6. Add adapter unit tests with synthetic fixtures. Certified sample XML is not checked into this repo.

Do not treat namespace prefixes, certifier envelope paths, or GCS download as domain behavior.