<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[中文OCR数据集分享]]></title><description><![CDATA[<p>分享一些中文OCR的数据集，其中包括单字、链接文本、屏幕截图、书本文字、低亮度、街景文本等很多场景形式的数据。方便大家使用不同的数据集进行索引。</p>
<ul>
<li>
<p><a href="https://ctwdataset.github.io" target="_blank" rel="noopener noreferrer nofollow">https://ctwdataset.github.io</a><br />
<img src="https://ctwdataset.github.io/static/img/gt_1044721_0_0_2048_2048.jpg" alt="" class="img-responsive img-markdown" /></p>
</li>
<li>
<p><a href="https://vision.cornell.edu/se3/coco-text-2/" target="_blank" rel="noopener noreferrer nofollow">https://vision.cornell.edu/se3/coco-text-2/</a><br />
<img src="/assets/uploads/files/1564371444328-5ef2fe17-1f2b-407b-b274-f230dbe0f96e-image.png" alt="5ef2fe17-1f2b-407b-b274-f230dbe0f96e-image.png" class="img-responsive img-markdown" /></p>
</li>
<li>
<p><a href="http://www.robots.ox.ac.uk/~vgg/data/scenetext/" target="_blank" rel="noopener noreferrer nofollow">http://www.robots.ox.ac.uk/~vgg/data/scenetext/</a><br />
这是一个算法，可以生成包含文本的图片</p>
</li>
<li>
<p><a href="https://github.com/wang-tf/Chinese_OCR_synthetic_data" target="_blank" rel="noopener noreferrer nofollow">https://github.com/wang-tf/Chinese_OCR_synthetic_data</a><br />
文本图片合成的中文版本</p>
</li>
<li>
<p><a href="http://www.robots.ox.ac.uk/~vgg/data/text/" target="_blank" rel="noopener noreferrer nofollow">http://www.robots.ox.ac.uk/~vgg/data/text/</a><br />
合成之后的英文和中文文本数据集</p>
</li>
<li>
<p><a href="https://pan.baidu.com/s/1dFda6R3#list/path=%2Fsharelink2436197375-565847762105022%2FSynthetic%20Chinese%20String%20Dataset&amp;parentPath=%2Fsharelink2436197375-565847762105022" target="_blank" rel="noopener noreferrer nofollow">https://pan.baidu.com/s/1dFda6R3#list/path=%2Fsharelink2436197375-565847762105022%2FSynthetic Chinese String Dataset&amp;parentPath=%2Fsharelink2436197375-565847762105022</a></p>
</li>
</ul>
<p>未完待续</p>
]]></description><link>http://t.manaai.cn/topic/165/中文ocr数据集分享</link><generator>RSS for Node</generator><lastBuildDate>Mon, 14 Sep 2026 11:22:09 GMT</lastBuildDate><atom:link href="http://t.manaai.cn/topic/165.rss" rel="self" type="application/rss+xml"/><pubDate>Mon, 29 Jul 2019 03:41:20 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to 中文OCR数据集分享 on Mon, 29 Jul 2019 04:28:44 GMT]]></title><description><![CDATA[<p>分享一些中文OCR的数据集，其中包括单字、链接文本、屏幕截图、书本文字、低亮度、街景文本等很多场景形式的数据。方便大家使用不同的数据集进行索引。</p>
<ul>
<li>
<p><a href="https://ctwdataset.github.io" target="_blank" rel="noopener noreferrer nofollow">https://ctwdataset.github.io</a><br />
<img src="https://ctwdataset.github.io/static/img/gt_1044721_0_0_2048_2048.jpg" alt="" class="img-responsive img-markdown" /></p>
</li>
<li>
<p><a href="https://vision.cornell.edu/se3/coco-text-2/" target="_blank" rel="noopener noreferrer nofollow">https://vision.cornell.edu/se3/coco-text-2/</a><br />
<img src="/assets/uploads/files/1564371444328-5ef2fe17-1f2b-407b-b274-f230dbe0f96e-image.png" alt="5ef2fe17-1f2b-407b-b274-f230dbe0f96e-image.png" class="img-responsive img-markdown" /></p>
</li>
<li>
<p><a href="http://www.robots.ox.ac.uk/~vgg/data/scenetext/" target="_blank" rel="noopener noreferrer nofollow">http://www.robots.ox.ac.uk/~vgg/data/scenetext/</a><br />
这是一个算法，可以生成包含文本的图片</p>
</li>
<li>
<p><a href="https://github.com/wang-tf/Chinese_OCR_synthetic_data" target="_blank" rel="noopener noreferrer nofollow">https://github.com/wang-tf/Chinese_OCR_synthetic_data</a><br />
文本图片合成的中文版本</p>
</li>
<li>
<p><a href="http://www.robots.ox.ac.uk/~vgg/data/text/" target="_blank" rel="noopener noreferrer nofollow">http://www.robots.ox.ac.uk/~vgg/data/text/</a><br />
合成之后的英文和中文文本数据集</p>
</li>
<li>
<p><a href="https://pan.baidu.com/s/1dFda6R3#list/path=%2Fsharelink2436197375-565847762105022%2FSynthetic%20Chinese%20String%20Dataset&amp;parentPath=%2Fsharelink2436197375-565847762105022" target="_blank" rel="noopener noreferrer nofollow">https://pan.baidu.com/s/1dFda6R3#list/path=%2Fsharelink2436197375-565847762105022%2FSynthetic Chinese String Dataset&amp;parentPath=%2Fsharelink2436197375-565847762105022</a></p>
</li>
</ul>
<p>未完待续</p>
]]></description><link>http://t.manaai.cn/post/364</link><guid isPermaLink="true">http://t.manaai.cn/post/364</guid><dc:creator><![CDATA[刘看山]]></dc:creator><pubDate>Mon, 29 Jul 2019 04:28:44 GMT</pubDate></item><item><title><![CDATA[Reply to 中文OCR数据集分享 on Tue, 30 Jul 2019 08:26:28 GMT]]></title><description><![CDATA[<p>在Github翻遍了也没有找到一个能够达到工业级别的解决方案。。。</p>
]]></description><link>http://t.manaai.cn/post/369</link><guid isPermaLink="true">http://t.manaai.cn/post/369</guid><dc:creator><![CDATA[WangXuanBT]]></dc:creator><pubDate>Tue, 30 Jul 2019 08:26:28 GMT</pubDate></item><item><title><![CDATA[Reply to 中文OCR数据集分享 on Wed, 31 Jul 2019 03:02:03 GMT]]></title><description><![CDATA[<p><a class="plugin-mentions-user plugin-mentions-a" href="http://t.manaai.cn/uid/188">@WangXuanBT</a> 工业级要多工业级？问题得拆分细化，不能笼统打死。<br />
目前OCR都是先检测，再做识别，没有端到端的。因为单单检测都比较复杂。<br />
检测可能也是最耗费时的。<br />
我们后续会出一个多字的版本，至于检测，可能现在更多的纠结在于精度和速度，以及一些对公式‘中英文混杂的特殊情况的处理。</p>
]]></description><link>http://t.manaai.cn/post/376</link><guid isPermaLink="true">http://t.manaai.cn/post/376</guid><dc:creator><![CDATA[刘看山]]></dc:creator><pubDate>Wed, 31 Jul 2019 03:02:03 GMT</pubDate></item></channel></rss>