ModelZoo
ModelZoo内模型数据定期更新,模型应用示例请参考ai-sdk
基础模型
- K1
- 推理引擎版本: spacemit-ort-2.0.6
- OS:bianbu-3.0
- date:2026-7-27
- K3
- 推理引擎版本: v2.0.6
- OS:bianbu-4.0rc3
- date:2026-7-27
测试方式
# 进入spacemit-ort库路径
# cd {spacemit_ort_lib}/
export LD_LIBRARY_PATH=./lib/
# 调整为自己的${model_path}(模型文件路径),${num of cores}(选择跑几个核心)
./bin/onnxruntime_perf_test ${model_path} -e spacemit -r 10 -x 1 -S 1 -s -c 1 -i "SPACEMIT_EP_INTRA_THREAD_NUM|${num of cores}" -I
# 输出信息如下
using SpaceMITExecutionProvider
setting SPACEMIT_EP_INTRA_THREAD_NUM : 4
Setting intra_op_num_threads to 1
Session creation time cost: 0.169475 s
First inference time cost: 109 ms
Total inference time cost: 0.0727021 s
Total inference requests: 10
Average inference time cost total: 7.270205 ms
Total inference run time: 0.0727619 s
Number of inferences per second: 137.435
Avg CPU usage: 62 %
Peak working set size: 91336704 bytes
Avg CPU usage:62
Peak working set size:91336704
Runs:10
Min Latency: 0.00720383 s
Max Latency: 0.00730163 s
P50 Latency: 0.00727787 s
P90 Latency: 0.00730163 s
P95 Latency: 0.00730163 s
P99 Latency: 0.00730163 s
P999 Latency: 0.00730163 s
# Average inference time cost total即单帧推理耗时
resnet
- K1
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms |
|---|---|---|---|---|---|
| resnet18 | int8 | 224x224 | 40.68 | 22.15 | 12.99 |
| resnet50 | int8 | 224x224 | 95.36 | 52.89 | 32.29 |
| resnet50 | fp16 | 224x224 | 674.48 | 348.92 | 213.96 |
- K3
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms | 8 Core/ms |
|---|---|---|---|---|---|---|
| resnet18 | int8 | 224x224 | 8.00 | 4.74 | 2.90 | 2.07 |
| resnet50 | int8 | 224x224 | 21.00 | 12.21 | 7.72 | 5.47 |
| resnet50.batch4 | int8 | 224x224 | 77.99 | 42.75 | 24.62 | 15.38 |
| resnet50 | fp16 | 224x224 | 37.64 | 21.95 | 14.89 | 11.36 |
mobilenet
- K1
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms |
|---|---|---|---|---|---|
| mobilenet_v1 | int8 | 224x224 | 28.99 | 15.17 | 9.14 |
| mobilenet_v2 | int8 | 224x224 | 28.83 | 17.86 | 11.81 |
| mobilenet_v3_small | fp16 | 224x224 | 25.82 | 15.52 | 9.99 |
| mobilenet_v3_large | fp16 | 224x224 | 65.79 | 39.79 | 25.28 |
- K3
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms | 8 Core/ms |
|---|---|---|---|---|---|---|
| mobilenet_v1 | int8 | 224x224 | 12.67 | 7.21 | 3.91 | 2.35 |
| mobilenet_v2 | int8 | 224x224 | 17.69 | 9.92 | 5.22 | 3.26 |
| mobilenet_v3_small | fp16 | 224x224 | 8.71 | 5.14 | 3.23 | 2.75 |
| mobilenet_v3_large | fp16 | 224x224 | 16.59 | 9.67 | 5.95 | 4.52 |
efficientnet
- K1
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms |
|---|---|---|---|---|---|
| efficientnet_v1_b0 | int8 | 224x224 | 80.79 | 46.10 | 29.91 |
| efficientnet_v1_b1 | int8 | 224x224 | 115.95 | 66.09 | 42.71 |
| efficientnet_v2_s | int8 | 224x224 | 162.11 | 91.75 | 57.08 |
| efficientnet_v1_b0 | fp16 | 224x224 | 136.75 | 79.19 | 53.62 |
| efficientnet_v1_b1 | fp16 | 224x224 | 193.29 | 112.98 | 78.21 |
| efficientnet_v2_s | fp16 | 224x224 | 597.90 | 323.13 | 185.05 |
- K3
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms | 8 Core/ms |
|---|---|---|---|---|---|---|
| efficientnet_v1_b0 | int8 | 224x224 | 40.84 | 22.51 | 12.89 | 9.85 |
| efficientnet_v1_b1 | int8 | 224x224 | 63.46 | 34.46 | 19.80 | 14.75 |
| efficientnet_v2_s | int8 | 224x224 | 58.41 | 33.46 | 19.80 | 13.53 |
| efficientnet_v1_b0 | fp16 | 224x224 | 43.35 | 23.67 | 14.30 | 10.91 |
| efficientnet_v1_b1 | fp16 | 224x224 | 63.37 | 35.19 | 20.96 | 15.63 |
| efficientnet_v2_s | fp16 | 224x224 | 82.89 | 47.68 | 29.77 | 20.67 |
vit
- K1
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms |
|---|---|---|---|---|---|
| vit_b_16 | int8 | 224x224 | 506.60 | 328.01 | 167.38 |
| vit_b_16 | fp16 | 224x224 | 2478.26 | 1398.33 | 765.65 |
- K3
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms | 8 Core/ms |
|---|---|---|---|---|---|---|
| vit_b_16 | int8 | 224x224 | 100.90 | 56.64 | 35.19 | 23.76 |
| vit_b_16 | fp16 | 224x224 | 138.46 | 86.37 | 62.11 | 48.75 |
yolov5
- K1
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms |
|---|---|---|---|---|---|
| yolov5n | int8 | 640x640 | 242.92 | 132.91 | 79.59 |
| yolov5s | int8 | 640x640 | 465.30 | 244.70 | 142.46 |
| yolov5m | int8 | 640x640 | 959.66 | 493.48 | 273.47 |
- K3
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms | 8 Core/ms |
|---|---|---|---|---|---|---|
| yolov5n | int8 | 640x640 | 43.76 | 24.11 | 14.36 | 9.57 |
| yolov5s | int8 | 640x640 | 73.40 | 40.20 | 23.95 | 15.81 |
| yolov5m | int8 | 640x640 | 152.03 | 81.90 | 45.88 | 28.97 |
yolov6
- K1
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms |
|---|---|---|---|---|---|
| yolov6n | int8 | 640x640 | 173.42 | 93.21 | 55.98 |
| yolov6s | int8 | 640x640 | 438.71 | 224.39 | 123.60 |
- K3
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms | 8 Core/ms |
|---|---|---|---|---|---|---|
| yolov6n | int8 | 640x640 | 31.90 | 18.04 | 10.95 | 7.53 |
| yolov6s | int8 | 640x640 | 65.37 | 35.98 | 21.33 | 13.42 |
yolov8
- K1
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms |
|---|---|---|---|---|---|
| yolov8n | int8 | 640x640 | 233.76 | 123.09 | 73.74 |
| yolov8s | int8 | 640x640 | 515.20 | 273.62 | 150.33 |
| yolov8m | int8 | 640x640 | 1050.21 | 529.48 | 298.47 |
- K3
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms | 8 Core/ms |
|---|---|---|---|---|---|---|
| yolov8n | int8 | 640x640 | 41.41 | 23.08 | 13.85 | 9.51 |
| yolov8s | int8 | 640x640 | 74.50 | 41.19 | 24.89 | 16.65 |
| yolov8m | int8 | 640x640 | 157.80 | 85.68 | 48.90 | 32.31 |
yolov8-seg
- K3
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms | 8 Core/ms |
|---|---|---|---|---|---|---|
| yolov8n-seg | int8 | 640x640 | 61.59 | 33.57 | 19.40 | 12.75 |
| yolov8s-seg | int8 | 640x640 | 103.27 | 56.33 | 33.24 | 21.51 |
| yolov8m-seg | int8 | 640x640 | 204.94 | 109.94 | 61.93 | 39.61 |
yolov8-pose
- K3
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms | 8 Core/ms |
|---|---|---|---|---|---|---|
| yolov8n-pose | int8 | 640x640 | 46.42 | 26.18 | 16.16 | 11.16 |
| yolov8s-pose | int8 | 640x640 | 81.61 | 45.33 | 27.87 | 18.64 |
| yolov8m-pose | int8 | 640x640 | 165.60 | 90.30 | 51.75 | 33.84 |
yolo12
- K1
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms |
|---|---|---|---|---|---|
| yolo12n | int8 | 640x640 | 365.20 | 194.36 | 118.23 |
| yolo12s | int8 | 640x640 | 795.61 | 443.95 | 246.50 |
| yolo12m | int8 | 640x640 | 1921.82 | 1004.42 | 581.78 |
- K3
| 模型名 | type | shape | 1 Core/ms | 2 Core/ms | 4 Core/ms | 8 Core/ms |
|---|---|---|---|---|---|---|
| yolo12n | int8 | 640x640 | 107.56 | 57.85 | 32.55 | 21.73 |
| yolo12s | int8 | 640x640 | 191.31 | 102.05 | 56.90 | 36.24 |
| yolo12m | int8 | 640x640 | 378.55 | 200.75 | 110.23 | 69.17 |
音频模型
- K1
| 模型名 | type | 4 Core/rtf |
|---|---|---|
| melotts | dyn_int8 | 0.984 |
| sensevoice | dyn_int8 | --- |
- K3
| 模型名 | type | 4 Core/rtf | 8 Core/rtf |
|---|---|---|---|
| melotts | dyn_int8 | 0.530 | --- |
| sensevoice | dyn_int8 | 0.1124 | 0.1380 |
大模型
- K3
- llama.cpp版本:0.1.1
- OS:bianbu-4.0rc3
- date:2026-5-26
测试方式
# 进入spacemit-llama.cpp库路径
# cd {spacemit-llama.cpp}/
export LD_LIBRARY_PATH=./lib/
# 调整为自己的${model_path}(模型文件路径),${num of cores}(选择跑几个核心)
./bin/llama-bench -m ${model_path} -t ${num of cores} -p 128 -n 128 -mmp 0 -fa 1 -ub 128
# 输出信息如下
CPU_RISCV64_SPACEMIT: tcm is available, blk_size: 393216, blk_num: 8, is_fake_tcm: 0
CPU_RISCV64_SPACEMIT: num_cores: 16, num_perfer_cores: 8, perfer_core_arch_id: a064, exclude_main_thread: 0, use_ime1: 0, use_ime2: 1, mem_backend: HPAGE, cpu_mask: ff00, aicpu_id_offset: 8
CPU_RISCV64_SPACEMIT: alloc_chunk: open(/dev/tcm_sync_mem) failed, errno=2
CPU_RISCV64_SPACEMIT: failed to allocate init_barrier from shared mem, falling back to heap
| model | size | params | backend | threads | n_ubatch | fa | mmap | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | ---: | --------------: | -------------------: |
| qwen3 0.6B Q4_0 | 358.78 MiB | 596.05 M | CPU | 8 | 128 | 1 | 0 | pp128 | 499.75 ± 0.22 |
| qwen3 0.6B Q4_0 | 358.78 MiB | 596.05 M | CPU | 8 | 128 | 1 | 0 | tg128 | 53.35 ± 0.03 |
Qwen
- K3
| 模型名 | 量化类型 | PP128 (token/s) | TG128 (token/s) | PP1280 (token/s) | TG1280 (token/s) |
|---|---|---|---|---|---|
| qwen3-0.6B | Q4_0 | 499.75 | 53.35 | - | - |
| qwen3-1.7B | Q4_0 | 229.79 | 23.11 | - | - |
| qwen3-4B | Q4_0 | 76.44 | 11.03 | - | - |
| qwen3-moe-30B-A3B | Q4_0 | 55.67 | 12.32 | 44.03 | 11.17 |
| qwen3.5-0.8B | Q4_0 | 182.69 | 29.33 | - | - |
| qwen3.5-2B | Q4_1 | 112.22 | 16.15 | - | - |
HunYuan
- K3
| 模型名 | 量化类型 | PP128 (token/s) | TG128 (token/s) | PP1280 (token/s) | TG1280 (token/s) |
|---|---|---|---|---|---|
| HY-MT1.5-1.8B | Q4_K_M | 157.81 | 20.15 | - | - |
Llama
- K3
| 模型名 | 量化类型 | PP128 (token/s) | TG128 (token/s) | PP1280 (token/s) | TG1280 (token/s) |
|---|---|---|---|---|---|
| llama2-7B | Q4_0 | 50.40 | 7.07 | - | - |
多模态大模型
- K3
测试方式
以qwen3vlencoder为例
export LD_LIBRARY_PATH=./spacemit-llama.cpp/lib:./spacemit_ort/lib
export SPACEMIT_EP_DENSE_ACCURACY_LEVEL=1
llama-server -m qwen3vl-30b-text-q4_1.gguf --media-backend smt --smt-config-dir ./ -ctk f16 -ctv f16 -t 8 -c 1024 --host 0.0.0.0 --port 8080 --reasoning-budget 0 --reasoning off
详细参数含义见llama.cpp.md
VLM
- K3
| 模型名 | 图像规格 | LLM 8 Core + VisionEncoder 4 Core/ms | LLM 8 Core + VisionEncoder 8 Core/ms |
|---|---|---|---|
| fastvlm-0.5B | 512*512 | 256.47 | 164.50 |
| Qwen3-VL-30B-A3B | 768*768 | 7928.13 | 4753.55 |
| Qwen3.5-0.8B | 384*384 | 340.42 | 245.61 |
| Qwen3.5-2B | 384*384 | 901.56 | 794.03 |
| Qwen3.5-4B | 384*384 | 904.73 | 798.71 |
ASR
- K3
| 模型名 | 量化方式 | 线程配置 | RTF |
|---|---|---|---|
| Qwen3-ASR 0.6B | Q4_0 + 动态量化 ONNX | LLM 8 / AudioEncoder 4 | 0.169 |
| Qwen3-ASR 1.7B | Q4_0 + 动态量化 ONNX | LLM 8 / AudioEncoder 4 | 0.329 |
| Fun-ASR Nano | Q4_K_M + 量化 ONNX | LLM 4 / AudioEncoder 4 | 0.247 |
| Gemma4 ASR E2B | Q4_0 + 量化 ONNX | llama-server 8 / AudioEncoder 默认 | 0.689(转写)/ 0.578(翻译) |
Qwen3-ASR 0.6B 和 1.7B 均使用 001_zh_daily_weather.wav 和
004_zh_selling_sausages.wav 连续测试 3 轮,总音频时长 47.331 秒,总处理时间分别为
7.977 秒和 15.572 秒。Fun-ASR Nano 使用中文、英文、日文三条音频测试,单轮总音频
时长 28.558 秒,预热后处理时间 7.063 秒。Gemma4 ASR 中文转写使用
004_zh_selling_sausages.wav,音频时长 14.158 秒,处理时间 9.758 秒;翻译使用
日语、韩语、粤语样例连续测试 2 轮,总音频时长 34.104 秒,预热后处理时间
19.703 秒。首个动态 ONNX encoder session 初始化请求未计入。表中 RTF 为端到端
结果。由于测试集和任务不同,各模型的 RTF 不宜直接横向比较识别或翻译质量。