Hugging Face Blog·· 2024-12-05
Google 发布 PaliGemma 2 视觉语言模型系列
Welcome PaliGemma 2 – New vision language models by Google
AI 导读
Google 发布 PaliGemma 2,采用 SigLIP 图像编码器与 Gemma 2 文本解码器,提供 3B、10B 和 28B 参数的预训练模型,并支持 224、448 和 896 像素输入分辨率。发布内容还包括基于 DOCCI 微调的 3B、10B 模型、开放模型仓库、Transformers 集成及微调脚本;模型按 Gemma 许可分发。
来源:Hugging Face Blog · huggingface.co