跳到正文
原文
Hugging Face Blog·· 3 小时前AI 评分59

Liquid AI 发布端侧多模态决策模型 d1-3B 与 d1-omni-600M

Multimodal open d1 decision models for the edge

AI 导读

Liquid AI 发布开放权重决策模型 d1-3B 与实验性 d1-omni-600M,二者不生成 token,而用一次前向完成决策。

正文 · AI 翻译

今天,我们发布 d1 决策模型系列中的两款开源决策模型:d1-3B 和 d1-omni-600M(实验性)。

  • Decision Index 0.2.1 上 10B 以下的最佳决策模型:d1-3B 得分为 48.57,领先于所有 4B 和 9B 模型以及 Decider 35B-A3B(47.11)。
  • 多模态:d1-3B 支持文本和图像,而 d1-omni-600M 支持文本和图像,或文本和音频
  • 快速:d1-3B 在 NVIDIA Jetson AGX Thor 上回答一个问题需 16 ms,在 Jetson AGX Orin 上需 26 ms,在 Jetson Orin Nano 上需 50ms

我们如何为边缘设备构建决策模型

这些开源 d1 决策模型基于我们的 Liquid Foundation Models(LFMs)构建。与我们的生成式模型不同,决策模型不生成 token,而是在单次前向传播中给出答案。

d1-3B 和 d1-omni-600M 基于两种截然不同的骨干网络训练而成:

  • d1-3B 基于我们最新的纯解码器 VLM LFM2.5-VL-3B 训练而成。它接受文本和图像作为输入。
  • d1-omni-600M 基于双向编码器 LFM2.5-Encoder-350M 训练而成。它增加了视觉和音频编码器,以处理全部三种模态。它接受文本和图像,或文本和音频作为输入。该模型目前处于早期研究发布阶段,并在持续开发中。

基准测试结果

我们在七个公开数据集上对 d1-3B 和 d1-omni-600M 进行了基准测试,涵盖阅读理解、毒性检测、意图分类、医学问答和跨语言理解。d1-3B 的平均得分为 82.9,为表中最高,并高于 Decider 4B。d1-omni-600M 得分为 78.4,以仅四分之一的参数量超越 Decider 2B(77.1)。

基准测试 d1-omni-600M d1-3B Decider 2B Decider 4B
SQuAD 2.0 74.0 83.3 67.7 76.0
Civil Comments 95.8 93.3 93.6 92.8
MASSIVE intent 86.1 86.9 81.1 88.3
PubMedQA 61.3 68.3 65.7 63.3
BoolQ 77.7 86.3 87.3 89.0
XNLI 74.7 85.6 85.0 88.6
PAWS-X 79.5 76.4 59.5 69.8
平均 78.4 82.9 77.1 81.1

我们验证了 d1-3B 在标准视觉基准测试中保留了其 LFM2.5-VL-3B 骨干的视觉能力,并且 d1-omni-600M 能够处理全部三种模态。我们不报告任何视觉或音频基准测试结果,因为 Decision Index v0.3 仅包含一个私有视觉划分,而音频决策基准测试目前仍是一个开放性问题。

速度

我们与 NVIDIA 合作,在 NVIDIA GeForce RTX 4090、NVIDIA Jetson AGX Thor、Jetson AGX Orin 64 GB 和 Jetson Orin Nano 上,基于 NVIDIA 技术栈评估了 d1-3B。由于 d1-omni-600M 处于早期研究发布阶段,我们在本次发布中不报告其任何速度数据。

边缘推理。d1-3B 在每台实测设备上回答单个问题的时间均低于 50 ms。三个问题所需时间仅为一个问题的 1.3 倍,AGX Thor 从 16 ms 增至 20 ms。

一个问题 3 个问题 3.4K-token 状态 384px 图像 64 个状态,打包
Apple M5 Pro 30 ms 41 ms 640 ms 62 ms 78 / s
Jetson AGX Thor 16 ms 20 ms 220 ms 35 ms 262 / s
Jetson AGX Orin 64 GB 26 ms 35 ms 560 ms 83 ms 110 / s
Jetson Orin Nano 50 ms 73 ms 1,640 ms 202 ms 38 / s

GPU 推理。在 GPU 上,d1-3B 在两个平台上回答一个问题的时间均低于 10 ms,处理一张 384px 图像的时间均低于 18 ms。

一个问题 3 个问题 3.4K-token 状态 384px 图像 64 个状态,打包
NVIDIA RTX 4090 8 ms 21 ms 102 ms 17 ms 475 / s
AMD MI325X 9 ms 14 ms 44 ms 18 ms 1,106 / s

如何使用开源 d1 决策模型

当你需要快速、结构化的决策(包括多模态输入)时,请选用 d1 决策模型。d1-3B 在同等规模下提供最高的决策质量,而 d1-omni-600M 则适用于对占用空间有要求的场景。

安装依赖项(需要 transformers>=5.14):

pip install "transformers>=5.14" torch torchvision pillow

这些模型自带代码,因此请使用 trust_remote_code=True 加载:

import io
import urllib.request

import torch
from PIL import Image
from transformers import AutoModel

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True,
                                  dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)

# Several named questions over one text state, answered in one pass
questions = {
    "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
                          "fraud": "Suspected unauthorised use"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["Can wait", "Today", "Blocking the customer now"]},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))

# An image as the whole state
url = "http://images.cocodataset.org/val2017/000000039769.jpg"  # two cats on a sofa
photo = Image.open(io.BytesIO(urllib.request.urlopen(url).read()))
print(model.system_one(None, {"cats": {"type": "choice", "instructions": "How many cats are there?",
                                       "criteria": {"one": "One", "two": "Two", "more": "Three or more"}}},
                       images=[photo]))

# Many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))

为简洁起见,我们仅提供 d1-3B 的示例。有关如何运行的说明,请参阅 d1-omni-600M 模型卡。

开始使用开放的 d1 决策模型

两款决策模型均为开放权重,今天即可在 Hugging Face 上获取:

我们迫不及待想看看你会构建出什么。

引用

如果你使用了这项工作,请引用发布博客:

@article{liquidAI2026opend1,
  author  = {Liquid AI},
  title   = {Open d1: Edge decision models for text, vision, and audio},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/open-d1},
}

来源:Hugging Face Blog · huggingface.co