Hugging Face
Models
Datasets
Spaces
Community
Docs
Enterprise
Pricing
Log In
Sign Up
nguyenvulebinh
/
AVSRCocktail
like
0
Automatic Speech Recognition
Transformers
Safetensors
PyTorch
nguyenvulebinh/AVYT
English
avhubert_avsr
audio-visual-speech-recognition
multimodal
speech-recognition
lip-reading
cocktail-party
noise-robust
av-hubert
transformer
audio
video
english
lrs2
voxceleb2
ctc
attention
beam-search
multi-speaker
noisy-speech
arxiv:
2506.02178
Model card
Files
Files and versions
xet
Community
Train
Deploy
Use this model
67bfcfe
AVSRCocktail
1.72 GB
1 contributor
History:
2 commits
nguyenvulebinh
Upload AVHubertAVSR
67bfcfe
verified
4 months ago
.gitattributes
Safe
1.52 kB
initial commit
4 months ago
README.md
Safe
5.17 kB
Upload AVHubertAVSR
4 months ago
config.json
Safe
4.44 kB
Upload AVHubertAVSR
4 months ago
model.safetensors
Safe
1.72 GB
xet
Upload AVHubertAVSR
4 months ago