Delivery and mobility / Ride-hailing and delivery

Grab: Extracting text and fields from Southeast Asian documents

Company
Grab
Country
Singapore
Adoption stage
Pilot
Source published
Date basis
The date the source was published. It can differ from the date adoption started.
How the source was checked
Read the full source text

The work problem

Grab has to extract information from documents such as identity cards, driving licences and registration certificates. This is the first step of customer verification. Southeast Asia has many languages and formats, which makes the task especially hard. Existing OCR wavered when the format changed, while commercial LLMs were weak at understanding Southeast Asian languages, produced errors and hallucinations, and were slow to respond.

Technology and data

Grab chose Qwen2-VL 2B as the base model. It is a size that allows full fine-tuning, it is good at Thai and Vietnamese, and it reads the original resolution as it is. Training material came from two directions. Sentences in Southeast Asian languages were pulled from Common Crawl and turned into synthetic images with varied fonts and backgrounds, covering Indonesian, Thai, Vietnamese and English. For real documents, labels were extracted with the in-house automatic labelling tool Documint and then checked again by people. The LoRA approach fell short outside the Latin script, so it was changed to two-stage training that updates all parameters. Next, the vision encoder of Qwen2-VL 2B and the language decoder of Qwen2.5 0.5B were joined to build a new model of about 1B.

Results

Grab said the fully fine-tuned 2B model raised accuracy on Thai documents by 70 percentage points and on Vietnamese by 40 percentage points against the baseline. The new model of about 1B was within 3 percentage points of the 2B model on accuracy for most document types. It explained that latency was better than the 2B model, existing OCR and external APIs. It wrote that with external APIs the P99 latency stretched to three or four times the P50, which did not suit large-scale use.

Limits and open questions

The source covers model development and evaluation. There is no sentence stating when it moved into live operation or how far approval is automated.

Sources

Compiled from public sources. These are not results from ATF Works customers.

Read original (opens in a new tab)