#

vision-language-pretraining

Here are 31 public repositories matching this topic...

unitaryai / VTC-dataset

dataset video-understanding video-text-retrieval vision-language-pretraining vision-language-dataset

Updated May 1, 2024
Python

BUAADreamer / CCRK

[KDD 2024] Improving the Consistency in Cross-Lingual Cross-Modal Retrieval with 1-to-K Contrastive Learning

retrieval wit cross-modal cross-lingual mscoco multi30k image-text-search cross-modal-retrieval xlm-roberta swin-transformer cross-lingual-retrieval image-text-retrieval vision-language-pretraining iglue xflickrco kdd2024

Updated Jul 18, 2024
Python

xmed-lab / FD-SOS

MICCAI 2024 Oral: Vision-Language Open-Set Detectors for Bone Fenestration and Dehiscence Detection from Intraoral Images

object-detection teeth vision-language object-detector open-set-object-detection vision-language-pretraining miccai2024

Updated Jul 30, 2024
Python

LooperXX / ManagerTower

Code for ACL 2023 Oral Paper: ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation Learning

vision-language multi-modal-learning vision-language-pretraining vision-language-learning

Updated Dec 12, 2023
Python

ahmdtaha / distributed_sigmoid_loss

Unofficial implementation for Sigmoid Loss for Language Image Pre-Training

python3 pytorch unsupervised-learning vision-and-language multimodal-deep-learning self-supervised-learning vision-language contrastive-learning distributed-data-parallel vision-transformer vision-language-pretraining

Updated Sep 26, 2023
Python

jaisidhsingh / LoRA-CLIP

Easy wrapper for inserting LoRA layers in CLIP.

lora multimodal multimodal-deep-learning image-text-matching parameter-efficient-tuning vision-language-pretraining low-rank-adaptation

Updated Jun 16, 2024
Python

YyzHarry / vlm-fairness

Demographic Bias of Vision-Language Foundation Models in Medical Imaging

medical-imaging fairness subpopulation algorithmic-fairness bias-mitigation ood-generalization foundation-models vision-language-pretraining vision-language-model

Updated Feb 23, 2024
Python

unitaryai / VTC

VTC: Improving Video-Text Retrieval with User Comments

comments video-understanding multimodal-deep-learning video-text-retrieval vision-language-transformer vision-language-pretraining

Updated Aug 9, 2024
Python

ChenDelong1999 / ITRA

A codebase for flexible and efficient Image Text Representation Alignment

computer-vision deep-learning pytorch multimodal-learning vision-language-pretraining

Updated Jun 20, 2023
Python

adarobustness / adaptation_robustness

Evaluate robustness of adaptation methods on large vision-language models

robustness adaptation parameter-efficient-tuning vision-language-pretraining

Updated Aug 23, 2023
Shell

omipan / svl_adapter

SVL-Adapter: Self-Supervised Adapter for Vision-Language Pretrained Models

self-supervised-learning vision-language-pretraining

Updated Jan 11, 2024
Python

yiren-jian / BLIText

[NeurIPS 2023] Bootstrapping Vision-Language Learning with Decoupled Language Pre-training

multimodal-deep-learning vision-language-transformer vision-language-pretraining

Updated Dec 5, 2023
Python

alinlab / b2t

Bias-to-Text: Debiasing Unknown Visual Biases through Language Interpretation

explainable-ai vision-language-pretraining bias-and-fairness

Updated May 21, 2023
Python

TencentARC / FLM

Accelerating Vision-Language Pretraining with Free Language Modeling (CVPR 2023)

language-modeling vision-language-pretraining

Updated May 15, 2023
Python

TXH-mercury / COSA

Codes and Models for COSA: Concatenated Sample Pretrained Vision-Language Foundation Model

video-captioning video-qa video-retrieval vision-language-pretraining video-language-pretrainng

Updated Aug 1, 2023
Python

Zoky-2020 / SGA

Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training Models. [ICCV 2023 Oral]

adversarial-attack vision-language-pretraining

Updated Sep 6, 2023
Python

megvii-research / protoclip

📍 Official pytorch implementation of paper "ProtoCLIP: Prototypical Contrastive Language Image Pretraining" (IEEE TNNLS)

self-supervised-learning contrastive-learning vision-language-pretraining

Updated Nov 8, 2023
Python

HieuPhan33 / CVPR2024_MAVL

Multi-Aspect Vision Language Pretraining - CVPR2024

zero-shot-classification vision-language-pretraining vision-language-model zero-shot-segmentation medical-vision-and-language-pretraining

Updated Aug 20, 2024
Python

marslanm / Multimodality-Representation-Learning

This repository provides a comprehensive collection of research papers focused on multimodal representation learning, all of which have been cited and discussed in the survey just accepted https://dl.acm.org/doi/abs/10.1145/3617833 .

cross-modal multimodal-deep-learning multimodal-datasets transformer-models multimodal-pre-trained-model vision-language-pretraining multimodal-applications multimodal-pretext

Updated Oct 19, 2023

jusiro / FLAIR

FLAIR: A Foundation LAnguage-Image model of the Retina for fundus image understanding.

medical-imaging fundus-image-analysis foundation-models vision-language-pretraining

Updated May 15, 2024
Python

Improve this page

Add a description, image, and links to the vision-language-pretraining topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the vision-language-pretraining topic, visit your repo's landing page and select "manage topics."