smartytrios/document_data_extractor download history
smartytrios/document_data_extractor is a dataset on the Hugging Face Hub. In the last 30 days it was downloaded 17 times (5 in the last 7 days), and 129 times in total. It ranks #417,311 among datasets by monthly downloads.
Dataset Title: OCR-to-JSON Information Extraction Project Overview This dataset is specifically designed for fine-tuning Large Language Models (LLMs) to perform structured data extraction from Optical Character Recognition (OCR) outputs. The primary objective is to convert raw,
Models trained on document_data_extractor
1 models list it as training data.
- smartytrios/docintel_ocr_llama_3_2_gguf 104 downloads in 30 days
Open smartytrios/document_data_extractor on Hugging Face
Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.