GitHub - PaddlePaddle/PaddleOCR at b1b04c71874b642fcb0c33666189f96ea102741a

Branches Tags

Name		Name	Last commit message	Last commit date
Latest commit History 4,655 Commits
.github/ISSUE_TEMPLATE		.github/ISSUE_TEMPLATE
PPOCRLabel		PPOCRLabel
StyleText		StyleText
applications		applications
benchmark		benchmark
configs		configs
deploy		deploy
doc		doc
notebook		notebook
ppocr		ppocr
ppstructure		ppstructure
test_tipc		test_tipc
tools		tools
.clang_format.hook		.clang_format.hook
.gitignore		.gitignore
.pre-commit-config.yaml		.pre-commit-config.yaml
.style.yapf		.style.yapf
LICENSE		LICENSE
MANIFEST.in		MANIFEST.in
README.md		README.md
README_ch.md		README_ch.md
__init__.py		__init__.py
paddleocr.py		paddleocr.py
requirements.txt		requirements.txt
setup.py		setup.py
train.sh		train.sh

Repository files navigation

English | 简体中文

Introduction

PaddleOCR aims to create multilingual, awesome, leading, and practical OCR tools that help users train better models and apply them into practice.

Recent updates

2021.12.21 OCR open source online course starts. The lesson starts at 8:30 every night and lasts for ten days. Free registration: https://aistudio.baidu.com/aistudio/course/introduce/25207
2021.12.21 release PaddleOCR v2.4, release 1 text detection algorithm (PSENet), 3 text recognition algorithms (NRTR、SEED、SAR), 1 key information extraction algorithm (SDMGR, tutorial) and 3 DocVQA algorithms (LayoutLM, LayoutLMv2, LayoutXLM, tutorial).
PaddleOCR R&D team would like to share the key points of PP-OCRv2, at 20:15 pm on September 8th, Course Address.
2021.9.7 release PaddleOCR v2.3, PP-OCRv2 is proposed. The inference speed of PP-OCRv2 is 220% higher than that of PP-OCR server in CPU device. The F-score of PP-OCRv2 is 7% higher than that of PP-OCR mobile.
2021.8.3 released PaddleOCR v2.2, add a new structured documents analysis toolkit, i.e., PP-Structure, support layout analysis and table recognition (One-key to export chart images to Excel files).
2021.4.8 release end-to-end text recognition algorithm PGNet which is published in AAAI 2021. Find tutorial here；release multi language recognition models, support more than 80 languages recognition; especically, the performance of English recognition model is Optimized.
more

Features

PaddleOCR support a variety of cutting-edge algorithms related to OCR, and developed industrial featured models/solution PP-OCR and PP-Structure on this basis, and get through the whole process of data production, model training, compression, inference and deployment.

It is recommended to start with the “quick experience” in the document tutorial

Quick Experience

Web online experience for the ultra-lightweight OCR: Online Experience
Mobile DEMO experience (based on EasyEdge and Paddle-Lite, supports iOS and Android systems): Sign in to the website to obtain the QR code for installing the App
One line of code quick use: Quick Start

E-book: Dive Into OCR

Dive Into OCR 📚

Community

Join us👬: Scan the QR code below with your Wechat, you can join the official technical discussion group. Looking forward to your participation.
Contribution🏅️: Contribution page contains various tools and applications developed by community developers using PaddleOCR, as well as the functions, optimized documents and codes contributed to PaddleOCR. It is an official honor wall for community developers and a broadcasting station to help publicize high-quality projects.
Regular Season🎁: The community regular season is a point competition for OCR developers, covering four types: documents, codes, models and applications. Awards are selected and awarded on a quarterly basis. Please refer to the link for more details.

PP-OCR Series Model List（Update on September 8th）

Model introduction	Model name	Recommended scene	Detection model	Direction classifier	Recognition model
Chinese and English ultra-lightweight PP-OCRv3 model（16.2M）	ch_PP-OCRv3_xx	Mobile & Server	inference model / trained model	inference model / trained model	inference model / trained model
English ultra-lightweight PP-OCRv3 model（13.4M）	en_PP-OCRv3_xx	Mobile & Server	inference model / trained model	inference model / trained model	inference model / trained model
Chinese and English ultra-lightweight PP-OCRv2 model（11.6M）	ch_PP-OCRv2_xx	Mobile & Server	inference model / trained model	inference model / trained model	inference model / trained model
Chinese and English ultra-lightweight PP-OCR model (9.4M)	ch_ppocr_mobile_v2.0_xx	Mobile & server	inference model / trained model	inference model / trained model	inference model / trained model
Chinese and English general PP-OCR model (143.4M)	ch_ppocr_server_v2.0_xx	Server	inference model / trained model	inference model / trained model	inference model / trained model

For more model downloads (including multiple languages), please refer to PP-OCR series model downloads.
For a new language request, please refer to Guideline for new language_requests.
For structural document analysis models, please refer to PP-Structure models.

Tutorials

Environment Preparation
Quick Start
PP-OCR 🔥
- Quick Start
- Model Zoo
- Model training
- Model Compression
- Inference and Deployment
PP-Structure 🔥
Academic algorithms
Data Annotation and Synthesis
Datasets
Code Structure
Visualization
Community
New language requests
FAQ
References
License

Visualization more

Chinese OCR model

English OCR model

Multilingual OCR model

Guideline for New Language Requests

If you want to request a new language support, a PR with 1 following files are needed：

In folder ppocr/utils/dict, it is necessary to submit the dict text to this path and name it with {language}_dict.txt that contains a list of all characters. Please see the format example from other files in that folder.

If your language has unique elements, please tell me in advance within any way, such as useful links, wikipedia and so on.

More details, please refer to Multilingual OCR Development Plan.

License

This project is released under Apache 2.0 license

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Introduction

Features

Quick Experience

E-book: Dive Into OCR

Community

PP-OCR Series Model List（Update on September 8th）

Tutorials

Visualization more

Guideline for New Language Requests

License

About

Releases 13

Packages

Used by 2.8k

Contributors 197

Languages

License

PaddlePaddle/PaddleOCR

Folders and files

Latest commit

History

Repository files navigation

Introduction

Features

Quick Experience

E-book: Dive Into OCR

Community

PP-OCR Series Model List（Update on September 8th）

Tutorials

Visualization more

Guideline for New Language Requests

License

About

Topics

Resources

License

Stars

Watchers

Forks

Releases 13

Packages 0

Used by 2.8k

Contributors 197

Languages

Packages