Agentic Intelligence
A closed-loop manipulation policy: a frozen agent proposes programs in a nine-primitive robot DSL, a JEPA-style latent model imagines their outcomes, and a learned verifier ranks them before the arm moves.
Next-generation understanding of 3D computer vision and physical AI. We build perception that estimates what changes an outcome rather than what a scene looks like, world models that let an agent imagine before it commits, and decision interfaces that either certify an action or say plainly why they cannot.
What is the smallest thing a machine has to understand about the physical world in order to act in it reliably? Each card opens; the populated ones link straight to the papers.
Our work is not a set of unrelated projects. It is one loop, and each direction contributes a stage of it. Pick a stage.
Search by title, author or venue; filter by kind or by research direction. A left rule marks the items the lab led. Click a figure to enlarge it.
An asterisk marks the corresponding author and a dagger marks equal contribution. Items marked in preparation or under review are listed so the record is complete; they are not claimed as accepted work.
Each project page carries the full result set as interactive figures, including the negative results.
A closed-loop manipulation policy: a frozen agent proposes programs in a nine-primitive robot DSL, a JEPA-style latent model imagines their outcomes, and a learned verifier ranks them before the arm moves.
A training-free affordance layer between a frozen coding agent and a frozen grasp detector. Function ranks only inside what the arm can actually execute, and that ordering is the whole contribution.
Perception as sufficiency, not accuracy. A calibrated belief over the minimal outcome-relevant statistic, with a distribution-free coverage certificate and an honest abstention when it cannot be met.
A capability-factored benchmark for robot learning. Counterfactually paired episodes, mid-episode interventions and a first-class world-model track, scored as a vector over eleven axes rather than a single success rate.
Geometry-aware spatial reasoning for vision-language-action models, accepted at BMVC 2026. Structured Gaussian representations give the policy something to attend to besides pixels.
A robust dataset and risk-weighted learning for hand filtering, accepted at BMVC 2026. Built around deployment risk rather than aggregate accuracy, because not every error costs the same.
Data-efficient 6D pose estimation from RGB-D, published in IEEE Access. Vision foundation features paired with compact 3D Gaussian representations, with a granted Korean patent behind it.
Md Selim Sarowar leads CVπ Lab. His research asks how much of the physical world a machine actually has to represent in order to act in it, and his answer has been consistent across perception, world models and decision-making: less than the field assumes, but a different part of it. An accurate pose is neither necessary nor sufficient; what matters is the statistic that changes the outcome and whether the embodiment can execute the action that follows from it.
He is an M.Sc. (Research) candidate in Electronics Engineering at Yeungnam University, South Korea, ranked first in his cohort with a 4.5/4.5 CGPA, and a Graduate Research Assistant at the Advanced Visual Intelligence Lab under Prof. Sungho Kim. He holds the Global Korea Scholarship from NIIED, and previously worked on convex optimization as a Research Assistant at National Tsing Hua University, Taiwan. His thesis, Towards Physical AI: A Unified Framework for Perception, World Models, and Agentic Intelligence, is the through-line of everything on this page.
His work has appeared at BMVC and an ICML workshop, and in IEEE Access, Scientific Reports and Artificial Intelligence Review, with four granted patents across Korea, the United Kingdom and India. He reviews for PLOS ONE, IEEE Access, BMVC and ICML workshops.
Two papers accepted at BMVC 2026 (CORE A): GaussVLA, geometry-aware spatial reasoning for VLA models, and HNH40K, a robust dataset and risk-weighted learning for hand filtering.
GST-VLA, structured Gaussian spatial tokens for 3D depth-aware VLA models, accepted at an ICML 2026 workshop.
C3G-VM6D, our data-efficient 6D pose estimation system, published in IEEE Access, and Fusion VLM-Gait published in Scientific Reports. Both SCIE-Q1.
Korean patent granted for a compact 3D-Gaussian-based 6D object pose estimation method for physical AI, the fourth granted patent across Korea, the UK and India.
M.Sc. (Research) begins at Yeungnam University after a year of Korean language and literature under the GKS programme.
Moved to South Korea on the Global Korea Scholarship for the Korean language and literature programme at Yeungnam University.
Joined the Wireless Communications & Signal Processing Lab at National Tsing Hua University, Taiwan, as a Research Assistant.
B.Tech completed with distinction at KIIT, India.
We are a young group, which means the directions are still being shaped and there is room to own one. Two of the seven are deliberately open.
If you want to work on any of the seven directions, write with a short note on what you have built and which direction you would take. Code you have written matters more than a transcript.
We collaborate with the Brain AI Lab at Kyungpook National University and with MBZUAI. Humanoid robot vision and AR/VR are the two directions where we are actively looking for partners with hardware and a problem worth solving.
Every claim on this site links to a project page carrying the full result set, including the results that did not work. We report a failed pre-registered hypothesis, a mechanism we withdrew, and a null in as much detail as the wins.
Either address reaches the PI directly. Click the icon to copy.
The lab is based in South Korea, hosted alongside the Advanced Visual Intelligence Lab at Yeungnam University, Gyeongsan. We work with collaborators in South Korea, the United Arab Emirates and Taiwan.
Code, preprints and the PI's personal page.