smartytrios/document_data_extractor download history

smartytrios/document_data_extractor is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 17 times (5 in the last 7 days), and 129 times in total. It ranks #417,311 among datasets by monthly downloads.

Dataset Title: OCR-to-JSON Information Extraction Project Overview This dataset is specifically designed for fine-tuning Large Language Models (LLMs) to perform structured data extraction from Optical Character Recognition (OCR) outputs. The primary objective is to convert raw,

Models trained on document_data_extractor

1 models list it as training data.

Open smartytrios/document_data_extractor on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.