oshizo/japanese-text-image-retrieval-train download history
oshizo/japanese-text-image-retrieval-train is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 1,063 times (47 in the last 7 days), and 2,926 times in total. It ranks #18,464 among datasets by monthly downloads.
shunk031/JDocQAのtrain splitに含まれるPDFデータを画像化し、NDLOCRでOCRしたテキストとペアにしたデータセットです。OCRは長い辺を1200pxにリサイズした画像に対して実施しました。OCR結果には、読み取りに失敗した際の文字列「〓」が含まれます。本データセットに含めている画像は、長い辺を896px、700px、588pxのいずれかにリサイズしています。どのサイズとするかは主にページに含まれる文字数で決めました。 query列は、OCR結果の文字列に対しQwen/Qwen2.5-14B-Instructで生成したものです。3つの質問を生成させ、ランダムに1
Open oshizo/japanese-text-image-retrieval-train on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.