renhehuang/formosa-vision-finegrained download history

renhehuang/formosa-vision-finegrained is an image to text dataset on the Hugging Face Hub. In the last 30 days it was downloaded 38 times (5 in the last 7 days), and 479 times in total. It ranks #227,732 among datasets by monthly downloads.

Formosa Vision Fine-grained (Expanded) Dataset Summary 此資料集以台灣在地文化與地景為核心,提供具細節的中文描述,並保留原始圖像。 擴充版本針對每張圖像生成更長、更密集的語義描述,以強化模型在細節理解上的表現。 Motivation 『資料合成』FLAIR 的核心在於訓練模型「聽得懂細節」。這意味著「長文本」越具體、包含越多方位詞 (左上角、紅色物體旁...),模型學到的局部特徵就越好。因為在此階段會透過大型多模態模型生成豐富且長的中文描述夠「碎唸」(包含大量方位、顏色、材質

Open renhehuang/formosa-vision-finegrained on Hugging Face

Sister project: Paper Pulse, the upvote history of every Hugging Face Daily Paper.