Youngtaek Oh

Contact: {firstname}.{lastname} [at] kaist.ac.kr

yto2.jpg

I am a Ph.D. student advised by Prof. In So Kweon and Prof. Junmo Kim at KAIST. Previously, I was a research intern at Sony Group Corporation and LG AI Research.

My research centers on advancing multimodal representations across vision, language, and beyond. Currently, I am focused on improving vision-language compositionality and developing omnimodal embeddings to enable more structured and comprehensive multimodal understanding.

News

Aug 21, 2026 Our Omnimodal Embeddings paper is accepted to EMNLP Findings 2026!
Jun 11, 2025 Selected as a Outstanding Reviewer at CVPR 2025!
Nov 5, 2024 Selected as a Top Reviewer at NeurIPS 2024!
Sep 20, 2024 Our VL compositionality paper is accepted to EMNLP 2024! See you in Miami🌊🏖️

Publications

  1. syn_omni.png
    Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings
    Youngtaek OhQiyu Wu, Hiromi WakakiJunmo KimYuki Mitsufuji
    Conference on Empirical Methods in Natural Language Processing (EMNLP Findings), 2026
  2. cap4bridge.png
    Cap4Bridge: Caption-Guided Cross-Modal Contextualization With Stochastic Augmentation for Text-Video Retrieval
    MinJu Jeon, Hyungee KimSi-Woo KimYoungtaek Oh, Soeun LeeDong-Jin Kim
    IEEE Access, 2026
  3. fsc-clip-teaser.png
    Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
    Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024
    Oral presentation
  4. vl_compo.png
    Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
    Youngtaek Oh, Pyunghwan Ahn, Jinhyung Kim, Gwangmo Song, Soonyoung LeeIn So KweonJunmo Kim
    CVPR Workshop on ‘What is Next in Multimodal Foundation Models?’ (MMFM), 2024
  5. 2023_cviu_badapters.png
    Empirical study on using Adapters for debiased Visual Question Answering
    Computer Vision and Image Understanding (CVIU), 2023
  6. nice_teaser.png
    NICE: CVPR 2023 challenge on zero-shot image captioning
    Taehoon Kim, Pyunghwan Ahn, Sangyun Kim, Sihaeng Lee, Mark Marsden, Alessandra Sala, Seung Hwan Kim, Honglak Lee, Kyounghoon Bae, and 33 more authors
    preprint, 2024
    2nd place in the Image Captioning Challenge at the NICE Workshop, CVPR 2023
  7. self_sufficient.png
    Self-Sufficient Framework for Continuous Sign Language Recognition
    International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2023
    Oral presentation, Top 3% recognition of all accepted papers
  8. signing_outside.png
    Signing Outside the Studio: Benchmarking Background Robustness for Continuous Sign Language Recognition
    British Machine Vision Conference (BMVC), 2022
  9. daso.png
    DASO: Distribution-Aware Semantics-Oriented Pseudo-Label for Imbalanced Semi-Supervised Learning
    Youngtaek OhDong-Jin KimIn So Kweon
    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022
  10. ksl_guide.png
    KSL-Guide: A Large-scale Korean Sign Language Dataset Including Interrogative Sentences for Guiding the Deaf and Hard-of-Hearing
    Soomin Ham, Kibaek ParkYoungjoon JangYoungtaek Oh, Seokmin Yun, Sukwon Yoon, Chang Jo KimHan-Mu ParkIn So Kweon
    International Conference on Automatic Face and Gesture Recognition (FG), 2021
  11. sideguide.png
    SideGuide: A Large-scale Sidewalk Dataset for Guiding Impaired People
    {Kibaek ParkYoungtaek Oh, Soomin HamKyungdon Joo}*, Hyokyoung Kim, Hyoyoung KumIn So Kweon (*: equally contributed)
    International Conference on Intelligent Robots and Systems (IROS), 2020

Honors and Awards

Jun 2025 Outstanding Reviewer, CVPR 2025  [Homepage]
Nov 2024 Top Reviewer, NeurIPS 2024  [Homepage]
Jun 2023 Top 3% Recognition Certificates, ICASSP 2023   [Certificate]
May 2023 2nd place in NICE Challenge ($5000), NICE Workshop at CVPR 2023   [Press] [Code]
Oct 2022 Outstanding Reviewer, ECCV 2022  [Homepage]