A curated list of promising OCR resources

2 years after


A curated list of promising OCR resources


OcrKing 源自2009年初 Aven 在数据挖掘中的自用项目,在对技术的执着和爱好的驱动下积累已近七载经多年的积累和迭代,如今已经进化为云架构的集多层神经网络与深度学习于一体的OCR识别系统2010年初为方便更多用户使用,特制作web版文字OCR识别,从始至今 OcrKing一直提供免费识别服务及开发接口,今后将继续提供免费云OCR识别服务。OcrKing从未做过推广,

但也确确实实默默地存在,因为他相信有需求的朋友肯定能找得到。欢迎把 OcrKing 在线识别介绍给您身边有类似需求的朋友!希望这个工具对你有用,谢谢各位的支持!

OcrKing 能做什么?

OcrKing 是一个免费的快速易用的在线云OCR平台,可以将PDF及图片中的内容识别出来,生成一个内容可编辑的文档。支持多种文件格式输入及输出,支持多语种(简体中文,繁体中文,英语,日语,韩语,德语,法语等)识别,支持多种识别方式, 支持多种系统平台, 支持多形式API调用!

* [Tesseract-OCR](           

* [Tesseract.js is a pure Javascript port of the popular Tesseract OCR engine. ](
* [ Ocular is a state-of-the-art historical OCR system. ]( 

* [sfhistory  Making a map of historical SF photos -博文4所带库 ](                              

* [ocropy-论文1所带库 by Adnan Ul-Hasan](              

* [ A small C++ implementation of LSTM networks, focused on Adnan Ul-Hasan ](

* [ End to end OCR system for Telugu. Based on Convolutional Neural Networks. ]( )    

* [ Telugu OCR framework using RNN, CTC in Theano & Python3. ](

* [ Recurrent Neural Network and Long Short Term Memory (LSTM) with Connectionist Temporal Classification implemented in Theano. Includes a Toy training example. ]( )

* [ implement CTC with keras? #383 ](         

* [mxnet and ocr ](        

* [ An OCR-system based on Torch using the technique of LSTM/GRU-RNN, CTC and referred to the works of rnnlib and clstm.](

* [ pure javascript lstm rnn implementation based on ocropus ](

* ['caffe-ocr - OCR with caffe deep learning framework' by pannous ](     

* [ A implementation of LSTM and CTC to recognize image without splitting ](

* [ RNNSharp is a toolkit of deep recurrent neural network which is widely used for many different kinds of tasks, such as sequence labeling. It's written by C# language and based on .NET framework 4.6 or above version. RNNSharp supports many different types of RNNs, such as BPTT and LSTM RNN, forward and bi-directional RNNs, and RNN-CRF. ](           

* [warp-ctc A fast parallel implementation of CTC, on both CPU and GPU. by  BAIDU](        

Connectionist Temporal Classification is a loss function useful for performing supervised learning on sequence data, without needing an alignment between input data and labels. For example, CTC can be used to train end-to-end systems for speech recognition, which is how we have been using it at Baidu's Silicon Valley AI Lab.

Warp-CTC是一个可以应用在CPU和GPU上高效并行的CTC代码库 (library) 介绍 CTCConnectionist Temporal Classification作为一个损失函数,用于在序列数据上进行监督式学习,不需要对齐输入数据及标签。比如,CTC可以被用来训练端对端的语音识别系统,这正是我们在百度硅谷试验室所使用的方法。 端到端 系统 语音识别

* [ Test mxnet with own trained model,用训练好的网络模型进行数字,少量汉字,特殊字符(./等)的识别(总共有210类)](           

* [ An expandable and scalable OCR pipeline ](         

* [OpenOCR makes it simple to host your own OCR REST API.](          

* [ OCRmyPDF   uses Tesseract for OCR, and relies on its language packs. ](       

* [ OwncloudOCR uses tesseract OCR and OCRmyPDF for reading text from images and images in PDF files. ](

* [ Nextcloud OCR (optical character recoginition) processing for images and PDF with tesseract-ocr, OCRmyPDF and php-native message queueing for asynchronous purpose. ](

* [ 多标签分类,端到端的中文车牌识别基于mxnet, End-to-End Chinese plate recognition base on mxnet](      
* [中国二代身份证光学识别 ](      
* [ SwiftOCR:Fast and simple OCR library written in Swift ](     
* [Attention-OCR :Visual Attention based OCR ](            
* [ Added support for CTC in both Theano and Tensorflow along with image OCR example. #3436](     
* [EasyPR是一个开源的中文车牌识别系统,其目标是成为一个简单、高效、准确的车牌识别库。](      
* [Deep Embedded Clustering  for OCR based on caffe](       
* [ Deep Embedded Clustering  for OCR based on  MXNet](     
* [ The minimum OCR server by Golang The minimum OCR server by Golang, and a tiny sample application of gosseract.](        
* [ A comparasion among different variant of gradient descent algorithm This script implements and visualizes the performance the following algorithms, based on the MNIST hand-written digit recognition dataset:](     
* [ A curated list of resources dedicated to scene text localization and recognition ](    
* [ Convolutional Recurrent Neural Network (CRNN) for image-based sequence recognition. ](   

* [ Implementation of the method proposed in the papers " TextProposals: a Text-specific Selective Search Algorithm for Word Spotting in the Wild" and "Object Proposals for Text Extraction in the Wild" (Gomez & Karatzas), 2016 and 2015 respectively. ](          

* [ Word Spotting and Recognition with Embedded Attributes ](          

* [ Part of eMOP: Franken+ tool for creating font training for Tesseract OCR engine from page images. ](     

* [NOCR NOCR is an open source C++ software package for text recognition in natural scenes, based on OpenCV. The package consists of a library, console program and GUI program for text recognition.](      

* [An OpenCV based OCR system, base to other projects Uses Histogram of Oriented Gradients (HOG) to extract characters features and Support Vector Machines as a classifier. It serves as basis for other projects that require OCR functionality.](    

* [Recognize bib numbers from racing photos](

* [Automatic License Plate Recognition library](        

* [汽车挡风玻璃VIN码识别](

* [](

* [Image Recognition for the Democracy Project with codes](    

* [Tools to be evaluated prior to integration into Newman](

* [Text Recognition in Natural Images in Python](

## Papers

* [论文1 can we build language-independent ocr using lstm networks by Adnan Ul-Hasan](                  

* [Adnan Ul-Hasan的博士论文](             
* [Applying OCR Technology for Receipt Recognition](     

* [An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition](

* [Reading Scene Text in Deep Convolutional Sequences](           

* [What You Get Is What You See:A Visual Markup Decompiler](
>     Building on recent advances in image caption generation and optical character recognition (OCR), we present a general-purpose, deep learning-based system to decompile an image into presentational markup. While this task is a well-studied problem in OCR, our method takes an inherently different, data-driven approach. Our model does not require any knowledge of the underlying markup language, and is simply trained end-to-end on real-world example data. The model employs a convolutional network for text and layout recognition in tandem with an attention-based neural machine translation system. To train and evaluate the model, we introduce a new dataset of real-world rendered mathematical expressions paired with LaTeX markup, as well as a synthetic dataset of web pages paired with HTML snippets. Experimental results show that the system is surprisingly effective at generating accurate markup for both datasets. While a standard domain-specific LaTeX OCR system achieves around 25% accuracy, our model reproduces the exact rendered image on 75% of examples. 

* [ Recursive Recurrent Nets with Attention Modeling for OCR in the Wild](
>   We present recursive recurrent neural networks with attention modeling (R2AM) for lexicon-free optical character recognition in natural scene images. The primary advantages of the proposed method are: (1) use of recursive convolutional neural networks (CNNs), which allow for parametrically efficient and effective image feature extraction; (2) an implicitly learned character-level language model, embodied in a recurrent neural network which avoids the need to use N-grams; and (3) the use of a soft-attention mechanism, allowing the model to selectively exploit image features in a coordinated way, and allowing for end-to-end training within a standard backpropagation framework. We validate our method with state-of-the-art performance on challenging benchmark datasets: Street View Text, IIIT5k, ICDAR and Synth90k.     

* [#ICML 2016#【通过DNN把数据空间映射到latent的特征空间做聚类,目标函数是最小化软分配与辅助分布直接的KL距离,来迭代优化,思想类似于t-SNE,只不过这里使用了DNN】《Unsupervised Deep Embedding for Clustering Analysis》](
> Clustering is central to many data-driven application domains and has been studied extensively in terms of distance functions and grouping algorithms.  Relatively little work has focused on learning  representations  for  clustering.   In  this paper,  we  propose  Deep  Embedded  Clustering (DEC), a method that simultaneously learns feature representations and cluster assignments using  deep  neural  networks.   DEC  learns  a  mapping from the data space to a lower-dimensional feature space in which it iteratively optimizes a
clustering  objective.   Our  experimental  evaluations on image and text corpora show significant improvement over state-of-the-art methods

## Blogs

* [Tesseract-OCR引擎入门](             

* OCR引擎Ocropus实战指南              
* [博文1 Training an Ocropus OCR model ](          
* [博文2  Extracting text from an image using Ocropus](   
* [博文3 Working with Ground Truth ](                         
* [博文4 Finding blocks of text in an image using Python, OpenCV and numpy](                

* [Applying OCR Technology for Receipt Recognition]( )         

* [Writing a Fuzzy Receipt Parser in Python](
* [Number plate recognition with Tensorflow](      
* [车牌识别中的不分割字符的端到端(End-to-End)识别](         
* [端到端的OCR:基于CNN的实现](
* [ 腾讯OCR—自动识别技术,探寻文字真实的容颜 ](    

* [验证码识别](     

## Presentations

* 学霸君archsubmit上的演讲提到了他们的ocr算法 使用cnn来识别中文  链接: 密码: p92n

## Projects

* [ Project Naptha :highlight, copy, search, edit and translate text in any image](      

## Commercial products

* [ABBYY](

作者:chenqin 链接: 来源:知乎 著作权归作者所有。商业转载请联系作者获得授权,非商业转载请注明出处。

1,识别率极高。我使用过现在的答案总结里提到的所有软件,但遇到下面这样的表格,除了ABBYY还能保持95%以上的识别率之外(包括秦皇岛三个字),其他所有的软件全部歇菜,数字认错也就罢了,中文也认不出。血泪的教训。 2,自由度高。可以在同一页面手动划分不同的区块,每一个区块也可以分别设置表格或文字;简体繁体英文数字。而此时大部分软件还只能对一个页面设置一种识别方案,要么表格,要么文字。 3,批量操作方便。对于版式雷同的年鉴,将一页的版式设计好,便可以应用到其他页,省去大量重复操作。 4,可以保持原有表格格式,省去二次编辑。跨页识别表格时,选择“识别为EXCEL”,ABBYY可以将表格连在一起,产出的是一整个excel文件,分析起来就方便多了。 5,包括梯形校正,歪斜校正之类的许多图片校正方式,即使扫描得歪了,或者因为书本太厚而导致靠近书脊的部分文字扭曲,都可以校正回来。

* IRIS               

真正能把中文OCR做得比较专业的,一共也没几家,国内2家,国外2家。国内是文通和汉王,国外是ABBYY和IRIS(台湾原来有2家丹青和蒙恬,这两年没什么动静了)。像大家提到的紫光OCR、CAJViewer、MS Office、清华OCR、包括慧视小灵鼠,这些都是文通的产品或者使用文通的识别引擎,尚书则是汉王的产品,和中晶扫描仪捆绑销售的。这两家的中文识别率都是非常不错的。而国外的2家,主要特点是西方语言的识别率很好,而且支持多种西欧语言,产品化程度也很高,不过中文方面速度和识别率还是有差距的,当然这两年人家也是在不断进步。Google的开源项目,至少在中文方面,和这些家相比,各项性能指标水平差距还蛮大的呢。

作者:张岩 链接: 来源:知乎 著作权归作者所有。商业转载请联系作者获得授权,非商业转载请注明出处。


## OCR Databases


* [pics for testing 测试用图片 ]() 链接: 密码: 3izf

Related Repositories



A curated list of awesome Python frameworks, libraries and software ...



A curated list of awesome Python frameworks, libraries and software ...



A curated list of awesome Python frameworks, libraries and software ...



A curated list of awesome Python frameworks, libraries and software ...



A curated list of awesome Python frameworks, libraries and software ...

Top Contributors