Skip to main content
Atlas automatically classifies and extracts structured data from a wide range of documents used in financial workflows. You do not need to declare a document type at upload time — Atlas detects it from the image content and returns the appropriate fields in the document_type-specific payload within ocr_data. Each document type returns a distinct set of extracted fields, all following the same { "value": ..., "confidence_score": ... } structure. For the full list of fields per document type, see Supported Documents.

Document categories

These documents establish individual identity and are used for KYC verification in lending applications.
The document_quality_flag field (values: original, xerox_copy, photo_of_photo) is currently returned for PAN cards. Use it to flag low-quality copies for manual review before making credit decisions.
These documents verify the legal registration and operational status of a business entity.
These documents capture transaction and financing details for loan underwriting.
These documents are used in property loan and mortgage underwriting. Many are Maharashtra-specific revenue and registration records.
These document types process images rather than scanned paperwork, and are used to verify physical asset condition and identity at the point of delivery.
These documents are used for insurance underwriting, policy issuance, claims processing, and verification of the insured person, vehicle, income, and medical history.
These documents contain medical and insurance-claim information used for healthcare verification, treatment workflows, and claims processing.

Type-specific OCR fields

Each document_type in the ocr_data array returns a different field schema. For example, an AADHAAR document returns name, address, aadhaar_number, and date_of_birth, while an INVOICE returns customer_name, dealer_address, total_invoice_amount, and so on. Your processing logic should branch on document_type to read the correct fields.

Extraction failures

If Atlas cannot extract data from a document, the data object will be empty or partially populated, and the error_code and error_reason fields will contain a machine-readable code and a human-readable explanation.
Common error codes include UNRECOGNISED (document type not supported) and INVALID_DOC (document could not be processed, for example due to image quality or multiple subjects in frame). For details on every extracted field per document type, see Supported Documents.