Saltar al contenido principal

Purchase Document Ingestion Troubleshooting

Operational runbook for apps/backend/src/purchase-document-ingestion/.

Use with Purchase Document Ingestion and Purchase Document Extractor Port.


Scope​

Primary codepaths:

  • application/upload-purchase-documents.use-case.ts
  • application/extract-purchase-document.use-case.ts
  • application/ingest-email-attachment.use-case.ts
  • application/post-purchase-document.use-case.ts
  • infrastructure/extractors/fel-xml-extractor.adapter.ts
  • infrastructure/purchase-document-storage.service.ts
  • infrastructure/jobs/*
  • interfaces/purchase-document-ingestion.controller.ts
  • interfaces/webhooks/purchase-document-email-intake.controller.ts

1) Symptom triage​

SymptomLikely layer
Upload 400 Unsupported file typeController extension allowlist (XML/PDF only)
Upload 400 content does not match declared typePDF MIME sniff (file-signature.service)
Upload succeeds but list never shows the invoiceReceiver NIT mismatch → row deleted during extraction
Status stuck on received / "Still processing…"Bull worker down, or persist error that used to escape to Redis failed while the row stayed received
Status unparsedPDF uploaded without XML — expected, not a crash
Status failed immediatelyMalformed XML, unrecognized DTE, or persist error
Email never creates a draftInbox token, SendGrid parse URL, or no XML/PDF attachment
POST :id/post 400 "Only extracted documents can be posted"Undismissed review flags (needs_review)
POST :id/post 400 "Accounts Payable payment method is not configured"Missing accountsPayable payment method seed
POST :id/post 400 "At least one inventory line is required"All lines expense/ignore, or none mapped
POST :id/post 409Already posted (IngestionAlreadyPostedError)
Retry extraction 400Status is not failed (e.g. unparsed)
Preview 404 for a file that uploadeddocumentLink missing that kind (xml / pdf)
Signed download URL 403Clock skew, or GCS_PRIVATE_DOCUMENTS_BUCKET IAM

2) Capture and extraction checks​

Expected upload flow:

  1. Authenticated POST /purchase-document-ingestions/upload?businessId= with multipart field files.
  2. Use case groups by filename stem, stores objects, inserts ingestionStatus=received.
  3. Bull job extract with jobId = ingestionId.
  4. Worker runs ExtractPurchaseDocumentUseCase (idempotent if status is no longer received).

When a draft never appears:

  1. Confirm the XML Receptor NIT matches business.taxId. A mismatch deletes the row by design.
  2. Confirm the file was not reported in data.duplicates (reason: file_hash).
  3. Inspect backend logs for Extracting purchase document {id}.
  4. If the row exists as received for more than 10 minutes, the stalled sweep should re-enqueue. After 24 hours it fails with La extracción no pudo completarse en el tiempo esperado.

When status is failed:

  • Parser/type errors will not change on retry unless the extractor code changed. POST :id/retry-extraction is the path after a parser fix (re-upload is blocked by file-hash / FEL-authorization pairing).
  • Retry removes any existing Bull job with that jobId first. If retry appears to no-op, the old failed job was likely still in the removeOnFail: 500 set.

3) Email intake checks​

Webhook:

POST /webhooks/sendgrid/inbound-purchase-documents
  1. Confirm SendGrid Inbound Parse posts to that path (not the outbound Event Webhook). There is no signature header on this product.
  2. Confirm To matches inbox-{32 hex}@ingest.flowpos.app from GET .../inbox-address.
  3. Unknown tokens return 200 with no draft — this is intentional.
  4. Plain-text emails with no XML/PDF are discarded.
  5. Controller always returns 200 on processing errors; check logs for Failed to process inbound purchase-document email.
  6. Unverified From domains still ingest; look for unverified_sender on the draft after extraction.

4) Posting checks​

POST /purchase-document-ingestions/:id/post creates purchase.status = draft. If stock did not move, the purchase has not been submitted yet — that is expected.

Before posting, verify:

  • ingestionStatus === extracted (dismiss flags first)
  • dteType is FACT or FCAM
  • supplierId and locationId are set (PATCH :id)
  • at least one line is inventory with a product (variant required when hasVariants)
  • currency resolves from currencyId, then business-owned code, then global code
  • payment method key accountsPayable exists (getPaymentMethodByKey)

NCRE / NDEB drafts cannot post (PostingUnsupportedDteError).


5) Storage checks​

Bucket: GCS_PRIVATE_DOCUMENTS_BUCKET (private, signed-URL only).

Keys: purchase-documents/{businessId}/{uuid}/{filename}.

  • Upload throws if the env is unset.
  • GET :id replaces keys with 15-minute signed URLs.
  • In-app preview must use GET :id/files/{xml|pdf} (authenticated bytes). Do not expose object keys to the PWA.

6) Queue and sweep​

PieceBehavior
Queue namepurchase-document-extraction
Job nameextract
Job idingestion UUID
WorkerExtractPurchaseDocumentProcessor
Sweep@Cron(EVERY_10_MINUTES) via RetryStalledExtractionsProcessor
ClaimclaimForRetry + FOR UPDATE SKIP LOCKED (multi-instance safe)

If Cloud Run has no worker processing this queue, every draft stays received until the 24-hour timeout.