Purchase Document Extractor Port
Source-backed guide for the intake extractor contract in apps/backend/src/purchase-document-ingestion/.
Use this page when adding a new source (PDF OCR, spreadsheet, certifier pull). For the module workflow and HTTP surface, see Purchase Document Ingestion.
Intent
Every extractor produces the same source-agnostic ExtractedPurchaseDocument. Downstream use cases (supplier match, review flags, line resolution, posting) never branch on XML vs PDF vs future formats.
That is the same pattern as FEL's ElectronicCertificationProvider and ecommerce's EcommerceProvider: domain shape in the port, HTTP/XML details in the adapter.
Hexagonal map
| Layer | Codepaths | Responsibility |
|---|---|---|
| Domain | domain/purchase-document-extractor.port.ts | Canonical extracted header/lines, PurchaseDocumentExtractorPort |
| Application | application/extract-purchase-document.use-case.ts | Load stored bytes, call adapter, persist lines/flags, match supplier |
| Infrastructure | infrastructure/extractors/fel-xml-extractor.adapter.ts, purchase-document-extractor.registry.ts | FEL DTE XML parse, MIME routing |
| Interfaces | none for extraction | Extraction is never on the request path |
Port contract
export interface PurchaseDocumentExtractorPort {
supports(mimeType: string, filename: string): boolean;
extract(input: {
buffer: Buffer;
filename: string;
businessTaxId: string;
}): ExtractedPurchaseDocument | Promise<ExtractedPurchaseDocument>;
}
businessTaxId is the ingesting business NIT (from business.taxId), not the supplier. It exists so adapters can enforce the receiver-NIT check (FR-007).
ExtractedPurchaseDocument.source is currently "fel_xml" | "pdf" | "spreadsheet" | "fel_pull". Only fel_xml is implemented. PDF-only files never call an extractor; the use case sets unparsed instead.
Runtime adapter
FelXmlExtractorAdapter implements the port today.
| Concern | Adapter behavior |
|---|---|
| MIME | application/xml, text/xml, or filename ending .xml |
| Parser | fast-xml-parser with removeNSPrefix: true so SAT:/DTE:/sat: prefixes do not matter |
| Timezone | Guatemala fixed UTC-6 (-06:00); no TZ database |
| DTE types | FACT, FCAM, NCRE, NDEB only; others throw UnsupportedDocumentTypeError |
| Receiver NIT | mismatch throws DocumentRecipientMismatchError — the use case deletes the draft |
| Arithmetic | quantity × unit price vs line total collected as validationIssues (flags, not hard fail) |
PurchaseDocumentExtractorRegistry exists to pick an adapter via supports(). Extraction currently calls FelXmlExtractorAdapter directly after confirming an XML documentLink exists. When a second adapter ships, route through the registry instead of adding source branches in the use case.
Adding an extractor
- Implement
PurchaseDocumentExtractorPortunderinfrastructure/extractors/. - Map provider-specific fields into
ExtractedPurchaseDocumentonly — do not persist certifier XML paths on the ingestion row from the adapter. - Register the adapter in
PurchaseDocumentExtractorRegistry. - Keep receiver-NIT, unsupported-type, and arithmetic rules as domain errors /
validationIssues, not HTTP status codes. - Do not run extraction in the controller. Enqueue
purchase-document-extraction/extractlike the XML path. - Add adapter unit tests with synthetic fixtures. Certified sample XML is not checked into this repo.
Do not treat namespace prefixes, certifier envelope paths, or GCS download as domain behavior.
Related docs
- Purchase Document Ingestion
- FEL Provider Port — outbound certification; this port is inbound supplier XML
- Purchase Document Ingestion Troubleshooting