FW-PaddleOCR-VL-1.6
Models available for use with Fireworks on Foundry deliver optimized, best-in-class performance on the Fireworks Inference Cloud. Fireworks on Foundry is a Non-Microsoft Product. The following terms apply to a Customer's use of Fireworks on Foundry: When you use Fireworks on Foundry, data is shared between Microsoft and Fireworks AI, Customer Data will be sent outside of Microsoft systems, Customer Data will not be processed pursuant to any Foundry data residency documentation, and different compliance and data handling rules will apply. See Trust Center - Fireworks AI for details. Customers are responsible for evaluating whether data sharing between Microsoft and Fireworks is appropriate for their organization's compliance requirements.
About this model
PaddleOCR-VL-1.6 is a compact approximately 0.9B-parameter vision-language model for document parsing. It supports image input and a 131K-token context window, and recognizes text, tables, formulas, charts, seals, and text locations. PaddleOCR-VL-1.6 introduces a region-aware data optimization framework that targets weak regions from the previous model and a progressive post-training recipe using curated data selection and reinforcement learning. It achieves a 96.33% score on OmniDocBench v1.6 and is architecture-compatible with PaddleOCR-VL-1.5 for plug-and-play migration.
Key model capabilities
- Optical character recognition and text spotting
- Table, formula, chart, and seal recognition
- Chinese ancient document and rare-character recognition
- 131K-token context window with image input
- State-of-the-art 96.33% score on OmniDocBench v1.6
- Compact approximately 0.9B-parameter architecture