mirror of
https://github.com/mudler/LocalAI.git
synced 2026-08-04 04:12:22 -04:00
gallery: add VibeVoice ASR BitNet variants (#11296)
Add the recommended TQ2 build and a smaller aggressive quantization for the CrispASR backend. Assisted-by: Codex:gpt-5 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
This commit is contained in:
committed by
GitHub
parent
0990be35b7
commit
896b4b6785
@@ -39854,6 +39854,142 @@
|
||||
- filename: vibevoice-asr-q4_k.gguf
|
||||
uri: huggingface://cstr/vibevoice-asr-GGUF/vibevoice-asr-q4_k.gguf
|
||||
sha256: f1e87bb5c25dd469b495759e59c4554c4e8ec254f36c5c659737ff3e61ace982
|
||||
- name: vibevoice-asr-bitnet-crispasr
|
||||
url: github:mudler/LocalAI/gallery/virtual.yaml@master
|
||||
variants:
|
||||
- model: vibevoice-asr-bitnet-embed-q8-crispasr
|
||||
- model: vibevoice-asr-bitnet-vae-q5-crispasr
|
||||
- model: vibevoice-asr-bitnet-both-q5-crispasr
|
||||
- model: vibevoice-asr-bitnet-vae-q4-crispasr
|
||||
- model: vibevoice-asr-bitnet-aggro-crispasr
|
||||
urls:
|
||||
- https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
|
||||
- https://huggingface.co/cstr/vibevoice-asr-bitnet-GGUF
|
||||
description: |
|
||||
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model for English, Chinese, French, Italian, Korean, Portuguese, and Vietnamese. Its BitNet-trained language-model projections use ternary TQ2_0 weights. This recommended CrispASR build keeps the VAE encoder in Q8_0 and the embedding in F16, and is approximately 1.55 GB.
|
||||
license: mit
|
||||
tags:
|
||||
- crispasr
|
||||
- asr
|
||||
- speech-recognition
|
||||
- stt
|
||||
- gguf
|
||||
- bitnet
|
||||
- multilingual
|
||||
overrides:
|
||||
backend: crispasr
|
||||
known_usecases:
|
||||
- transcript
|
||||
name: vibevoice-asr-bitnet-crispasr
|
||||
parameters:
|
||||
model: vibevoice-asr-bitnet-tq2.gguf
|
||||
files:
|
||||
- filename: vibevoice-asr-bitnet-tq2.gguf
|
||||
uri: huggingface://cstr/vibevoice-asr-bitnet-GGUF/vibevoice-asr-bitnet-tq2.gguf
|
||||
sha256: 0a3053dfeffc8675673de892caa52b39ea7558091b9b422ebaeba94862a4d4a2
|
||||
- name: vibevoice-asr-bitnet-embed-q8-crispasr
|
||||
url: github:mudler/LocalAI/gallery/virtual.yaml@master
|
||||
urls:
|
||||
- https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
|
||||
- https://huggingface.co/cstr/vibevoice-asr-bitnet-GGUF
|
||||
description: |
|
||||
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model. This approximately 1.34 GB CrispASR build uses ternary TQ2_0 language-model projections, a Q8_0 VAE encoder, and a Q8_0 embedding.
|
||||
license: mit
|
||||
tags: [crispasr, asr, speech-recognition, stt, gguf, bitnet, multilingual]
|
||||
overrides:
|
||||
backend: crispasr
|
||||
known_usecases: [transcript]
|
||||
name: vibevoice-asr-bitnet-embed-q8-crispasr
|
||||
parameters:
|
||||
model: vibevoice-asr-bitnet-embed-q8.gguf
|
||||
files:
|
||||
- filename: vibevoice-asr-bitnet-embed-q8.gguf
|
||||
uri: huggingface://cstr/vibevoice-asr-bitnet-GGUF/vibevoice-asr-bitnet-embed-q8.gguf
|
||||
sha256: 450e5ae0a66a7bbe7318487a751524e4d73da799316ccffbbef0d03281bec849
|
||||
- name: vibevoice-asr-bitnet-vae-q5-crispasr
|
||||
url: github:mudler/LocalAI/gallery/virtual.yaml@master
|
||||
urls:
|
||||
- https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
|
||||
- https://huggingface.co/cstr/vibevoice-asr-bitnet-GGUF
|
||||
description: |
|
||||
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model. This approximately 1.33 GB CrispASR build uses ternary TQ2_0 language-model projections, a Q5_0 VAE encoder, and an F16 embedding.
|
||||
license: mit
|
||||
tags: [crispasr, asr, speech-recognition, stt, gguf, bitnet, multilingual]
|
||||
overrides:
|
||||
backend: crispasr
|
||||
known_usecases: [transcript]
|
||||
name: vibevoice-asr-bitnet-vae-q5-crispasr
|
||||
parameters:
|
||||
model: vibevoice-asr-bitnet-vae-q5.gguf
|
||||
files:
|
||||
- filename: vibevoice-asr-bitnet-vae-q5.gguf
|
||||
uri: huggingface://cstr/vibevoice-asr-bitnet-GGUF/vibevoice-asr-bitnet-vae-q5.gguf
|
||||
sha256: f31ec8b858ba69a425c626373e65f17bf73e383efe9a1e125f18196c6495c57f
|
||||
- name: vibevoice-asr-bitnet-both-q5-crispasr
|
||||
url: github:mudler/LocalAI/gallery/virtual.yaml@master
|
||||
urls:
|
||||
- https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
|
||||
- https://huggingface.co/cstr/vibevoice-asr-bitnet-GGUF
|
||||
description: |
|
||||
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model. This approximately 1.12 GB CrispASR build uses ternary TQ2_0 language-model projections, a Q5_0 VAE encoder, and a Q8_0 embedding.
|
||||
license: mit
|
||||
tags: [crispasr, asr, speech-recognition, stt, gguf, bitnet, multilingual]
|
||||
overrides:
|
||||
backend: crispasr
|
||||
known_usecases: [transcript]
|
||||
name: vibevoice-asr-bitnet-both-q5-crispasr
|
||||
parameters:
|
||||
model: vibevoice-asr-bitnet-both-q5.gguf
|
||||
files:
|
||||
- filename: vibevoice-asr-bitnet-both-q5.gguf
|
||||
uri: huggingface://cstr/vibevoice-asr-bitnet-GGUF/vibevoice-asr-bitnet-both-q5.gguf
|
||||
sha256: 3ad6552a214dab12786c3ae34afbbabef5701059a56a6e77cfeaf00c7cf02b80
|
||||
- name: vibevoice-asr-bitnet-vae-q4-crispasr
|
||||
url: github:mudler/LocalAI/gallery/virtual.yaml@master
|
||||
urls:
|
||||
- https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
|
||||
- https://huggingface.co/cstr/vibevoice-asr-bitnet-GGUF
|
||||
description: |
|
||||
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model. This approximately 1.26 GB CrispASR build uses ternary TQ2_0 language-model projections, a Q4_0 VAE encoder, and an F16 embedding.
|
||||
license: mit
|
||||
tags: [crispasr, asr, speech-recognition, stt, gguf, bitnet, multilingual]
|
||||
overrides:
|
||||
backend: crispasr
|
||||
known_usecases: [transcript]
|
||||
name: vibevoice-asr-bitnet-vae-q4-crispasr
|
||||
parameters:
|
||||
model: vibevoice-asr-bitnet-vae-q4.gguf
|
||||
files:
|
||||
- filename: vibevoice-asr-bitnet-vae-q4.gguf
|
||||
uri: huggingface://cstr/vibevoice-asr-bitnet-GGUF/vibevoice-asr-bitnet-vae-q4.gguf
|
||||
sha256: 72be9d760b2edb25faa8fa8a75a29d2de76837803599f3373496542ce955367b
|
||||
- name: vibevoice-asr-bitnet-aggro-crispasr
|
||||
url: github:mudler/LocalAI/gallery/virtual.yaml@master
|
||||
urls:
|
||||
- https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
|
||||
- https://huggingface.co/cstr/vibevoice-asr-bitnet-GGUF
|
||||
description: |
|
||||
Microsoft VibeVoice-ASR-BitNet is a 1.5B-parameter multilingual speech recognition model for English, Chinese, French, Italian, Korean, Portuguese, and Vietnamese. This compact CrispASR build combines ternary TQ2_0 language-model projections with a Q4_0 VAE encoder and Q8_0 embedding, reducing the model to approximately 1.05 GB while producing the same JFK benchmark transcription as the larger builds published alongside it.
|
||||
license: mit
|
||||
tags:
|
||||
- crispasr
|
||||
- asr
|
||||
- speech-recognition
|
||||
- stt
|
||||
- gguf
|
||||
- bitnet
|
||||
- multilingual
|
||||
overrides:
|
||||
backend: crispasr
|
||||
known_usecases:
|
||||
- transcript
|
||||
name: vibevoice-asr-bitnet-aggro-crispasr
|
||||
parameters:
|
||||
model: vibevoice-asr-bitnet-aggro.gguf
|
||||
files:
|
||||
- filename: vibevoice-asr-bitnet-aggro.gguf
|
||||
uri: huggingface://cstr/vibevoice-asr-bitnet-GGUF/vibevoice-asr-bitnet-aggro.gguf
|
||||
sha256: 5f118693e984e70e33e8818921ea41b2997eec80866337610e3d14bdd5bce6d1
|
||||
- name: vibevoice-tts-crispasr
|
||||
url: github:mudler/LocalAI/gallery/virtual.yaml@master
|
||||
urls:
|
||||
|
||||
Reference in New Issue
Block a user