Computer Vision and Physical AI Lab

CVπ Lab

Next-generation understanding of 3D computer vision and physical AI. We build perception that estimates what changes an outcome rather than what a scene looks like, world models that let an agent imagine before it commits, and decision interfaces that either certify an action or say plainly why they cannot.

3D Computer Vision World Models Vision-Language-Action Robot Foundation Models Humanoid Robot Vision Medical AI AR / VR 3D Human Pose Estimation
0
Papers, preprints and patents on record
0
Papers at CORE A* / A venues
0
Granted patents across three jurisdictions
0
Venues served as a peer reviewer
Research directions

Seven directions, one question

What is the smallest thing a machine has to understand about the physical world in order to act in it reliably? Each card opens; the populated ones link straight to the papers.

How the directions compose

Perceive, imagine, decide, certify

Our work is not a set of unrelated projects. It is one loop, and each direction contributes a stage of it. Pick a stage.

Bibliography

Publications, preprints and patents

Search by title, author or venue; filter by kind or by research direction. A left rule marks the items the lab led. Click a figure to enlarge it.

An asterisk marks the corresponding author and a dagger marks equal contribution. Items marked in preparation or under review are listed so the record is complete; they are not claimed as accepted work.

Open work

Project pages and code

Each project page carries the full result set as interactive figures, including the negative results.

The Agentic Intelligence framework: propose, imagine, verify, execute

Agentic Intelligence

A closed-loop manipulation policy: a frozen agent proposes programs in a nine-primitive robot DSL, a JEPA-style latent model imagines their outcomes, and a learned verifier ranks them before the arm moves.

The AFFORD-X framework: gate grasp candidates on feasibility, then rank by function

AFFORD-X

A training-free affordance layer between a frozen coding agent and a frozen grasp detector. Function ranks only inside what the arm can actually execute, and that ordering is the whole contribution.

The MSP framework: belief encoder, outcome head, and a certified inference block

MSP Framework

Perception as sufficiency, not accuracy. A calibrated belief over the minimal outcome-relevant statistic, with a distribution-free coverage certificate and an honest abstention when it cannot be met.

The PRAXIS capability radar: an oracle filling every axis while learned baselines sit at zero on spatial and planning

PRAXIS

A capability-factored benchmark for robot learning. Counterfactually paired episodes, mid-episode interventions and a first-class world-model track, scored as a vector over eleven axes rather than a single success rate.

The GaussVLA model: geometry-aware spatial reasoning for vision-language-action

GaussVLA

Geometry-aware spatial reasoning for vision-language-action models, accepted at BMVC 2026. Structured Gaussian representations give the policy something to attend to besides pixels.

HNH40K: a dataset and risk-weighted learning for hand filtering

HNH40K

A robust dataset and risk-weighted learning for hand filtering, accepted at BMVC 2026. Built around deployment risk rather than aggregate accuracy, because not every error costs the same.

C3G-VM6D architecture for data-efficient 6D pose estimation

C3G-VM6D

Data-efficient 6D pose estimation from RGB-D, published in IEEE Access. Vision foundation features paired with compact 3D Gaussian representations, with a granted Korean patent behind it.

People

Principal Investigator

Md Selim Sarowar leads CVπ Lab. His research asks how much of the physical world a machine actually has to represent in order to act in it, and his answer has been consistent across perception, world models and decision-making: less than the field assumes, but a different part of it. An accurate pose is neither necessary nor sufficient; what matters is the statistic that changes the outcome and whether the embodiment can execute the action that follows from it.

He is an M.Sc. (Research) candidate in Electronics Engineering at Yeungnam University, South Korea, ranked first in his cohort with a 4.5/4.5 CGPA, and a Graduate Research Assistant at the Advanced Visual Intelligence Lab under Prof. Sungho Kim. He holds the Global Korea Scholarship from NIIED, and previously worked on convex optimization as a Research Assistant at National Tsing Hua University, Taiwan. His thesis, Towards Physical AI: A Unified Framework for Perception, World Models, and Agentic Intelligence, is the through-line of everything on this page.

His work has appeared at BMVC and an ICML workshop, and in IEEE Access, Scientific Reports and Artificial Intelligence Review, with four granted patents across Korea, the United Kingdom and India. He reviews for PLOS ONE, IEEE Access, BMVC and ICML workshops.

Education

M.Sc. (Research), Electronics Engineering Yeungnam University, South Korea · AI & Computer Vision track Sep 2024 – Aug 2027 · CGPA 4.5/4.5, ranked 1st · GKS Scholar
B.Tech, Electrical & Electronics Engineering Kalinga Institute of Industrial Technology, India Jul 2019 – Jul 2023 · CGPA 8.65/10 · Study in India Scholar

Research positions

Graduate Research Assistant Advanced Visual Intelligence Lab, Yeungnam University Sep 2024 – present · Advisor: Prof. Sungho Kim
Research Assistant Wireless Communications & Signal Processing Lab, National Tsing Hua University, Taiwan Sep 2023 – Aug 2024 · Convex optimization
Machine Learning Intern Diginique Tech Lab, IIT Roorkee, India Jul 2022 – Aug 2022

Fellowships & honours

Global Korea Scholarship (GKS/KGSP) NIIED, Republic of Korea Sep 2024 – Aug 2027 · full funding
NTHU International Graduate Students Assistantship National Tsing Hua University, Taiwan
Bangladesh Sweden Trust Fund travel grant Research internship in Taiwan, 2024
Honorary Young Diplomat Jeju Special Governance Province, South Korea

Peer review

PLOS ONESCIE-Q1 journal · Public Library of Science2026 – present
BMVC 202637th British Machine Vision Conference2026
ICML 2026 WorkshopInternational Conference on Machine Learning2026
IEEE AccessSCIE-Q1 journal · IEEE2025 – present

Teaching

Computer VisionTeaching Assistant · Yeungnam UniversityFall 2025
EE203000 Linear AlgebraTeaching Assistant · National Tsing Hua UniversitySpring 2024
English Language Teaching (ELTA)Teaching Assistant · Ministry of Education, TaiwanNov 2023 – Aug 2024

Collaborations

Brain AI LabKyungpook National University, South KoreaJoint research
MBZUAIMohamed bin Zayed University of Artificial Intelligence, UAEJoint research
Advanced Visual Intelligence LabYeungnam University, South KoreaHost lab
Updates

Lab news

July 2026

Two papers accepted at BMVC 2026 (CORE A): GaussVLA, geometry-aware spatial reasoning for VLA models, and HNH40K, a robust dataset and risk-weighted learning for hand filtering.

June 2026

GST-VLA, structured Gaussian spatial tokens for 3D depth-aware VLA models, accepted at an ICML 2026 workshop.

2026

C3G-VM6D, our data-efficient 6D pose estimation system, published in IEEE Access, and Fusion VLM-Gait published in Scientific Reports. Both SCIE-Q1.

2026

Korean patent granted for a compact 3D-Gaussian-based 6D object pose estimation method for physical AI, the fourth granted patent across Korea, the UK and India.

September 2025

M.Sc. (Research) begins at Yeungnam University after a year of Korean language and literature under the GKS programme.

August 2024

Moved to South Korea on the Global Korea Scholarship for the Korean language and literature programme at Yeungnam University.

September 2023

Joined the Wireless Communications & Signal Processing Lab at National Tsing Hua University, Taiwan, as a Research Assistant.

July 2023

B.Tech completed with distinction at KIIT, India.

Openings

Join the lab

We are a young group, which means the directions are still being shaped and there is room to own one. Two of the seven are deliberately open.

Students and interns

If you want to work on any of the seven directions, write with a short note on what you have built and which direction you would take. Code you have written matters more than a transcript.

  • Strong Python and at least one of PyTorch, JAX or C++
  • Comfort reading a paper and reimplementing its core claim
  • A willingness to report the negative result

Collaborators

We collaborate with the Brain AI Lab at Kyungpook National University and with MBZUAI. Humanoid robot vision and AR/VR are the two directions where we are actively looking for partners with hardware and a problem worth solving.

  • Humanoid and bimanual platforms
  • Clinical partners for explainable motion analysis
  • Mixed-reality and real-time reconstruction

How we work

Every claim on this site links to a project page carrying the full result set, including the results that did not work. We report a failed pre-registered hypothesis, a mechanism we withdrew, and a null in as much detail as the wins.

  • Reproducible: numbers recomputed from run logs at build time
  • Honest: nulls and limitations reported, not summarised away
  • Open: code and project pages released with the paper
Get in touch

Contact

Email

Either address reaches the PI directly. Click the icon to copy.

Where we are

The lab is based in South Korea, hosted alongside the Advanced Visual Intelligence Lab at Yeungnam University, Gyeongsan. We work with collaborators in South Korea, the United Arab Emirates and Taiwan.

Elsewhere

Code, preprints and the PI's personal page.