Experiments in OpenCLIP (main branch) with `timm` NaFlexViT audio encoder + modern text encoder for variable-time , variable-length text CLAP models
-
rwightman/naflexclap_base_pf8_pt16_moderntextp.160m
Zero-Shot Image Classification • Updated -
rwightman/naflexclap_base_pf8_pt16_moderntext.160m
Zero-Shot Image Classification • Updated -
rwightman/naflexclap_base_pf8_pt20_moderntextp.240m
Zero-Shot Image Classification • Updated -
rwightman/naflexclap_mediumd_pf4_pt20_moderntextp.160m
Zero-Shot Image Classification • Updated