Physical AI & Agentic Robotics
Embodied agents that plan and act in latent world models for robot manipulation, bridging perception, reasoning, and control toward physical AGI.
Open to research collaborations & PhD opportunities Get in touch
I am pursuing an M.Sc. in Computer Vision & AI at Yeungnam University (YU), where I am maintaining a perfect CGPA of 4.5/4.5 (Ranked 1st / Top 1). I am a Global Korea Scholarship (GKS) scholar funded by NIIED, with additional support from the RLRC & RISE industry-collaboration projects. I am a member of the Advanced Visual Intelligence Lab, supervised by Prof. Sungho Kim.
My research focuses on 3D computer vision & spatial intelligence, 6D pose estimation, vision-language-action (VLA) models, world models, and agentic robotics for robot manipulation. I aim to build embodied agents capable of perceiving, reasoning, and acting within the physical world, working toward physical AGI. My contributions in these areas include first-author publications at BMVC'26 and in IEEE Access (SCIE-Q1).
Before joining YU, I completed a one-year Korean Language & Literature program at KLI under the GKS program, and spent a year as a full-time Research Assistant in the WCSP Lab at National Tsing Hua University, Taiwan. I earned my B.Tech (Electronics & Electrical Engineering) as an SII Scholar from the Kalinga Institute of Industrial Technology (KIIT), and interned at IIT Roorkee on image processing for biomedical signals.
| Rank | University | Country / Territory |
|---|---|---|
| No university matches that filter. | ||
| Loading the THE 2026 top 400… | ||
THE publishes individual positions to 200, then equal bands (201-250, 251-300, 301-350, 351-400). As published by THE 2026 — reproduced for reference; rankings remain the property of their publisher.
| Rank | University | Country / Territory |
|---|---|---|
| No university matches that filter. | ||
| Loading the QS 2027 top 400… | ||
QS publishes individual positions to 700; "=" marks a shared rank. As published by QS 2027 — reproduced for reference; rankings remain the property of their publisher.
Embodied agents that plan and act in latent world models for robot manipulation, bridging perception, reasoning, and control toward physical AGI.
Recovering full object pose and geometry from images and point clouds for spatially-grounded, robust scene understanding.
Learning transferable visual representations, from image denoising and autoencoders to point-cloud understanding for downstream 3D tasks.
@inproceedings{Sarowar_2026_BMVC,
author = {Md Selim Sarowar and Md Tanvir Islam and Sungho Kim and Sangtae Ahn},
title = {GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model},
booktitle = {37th British Machine Vision Conference 2026, {BMVC} 2026, Lancaster, UK, November 23-26, 2026},
publisher = {BMVA},
year = {2026},
url = {https://bmva-archive.org.uk/bmvc/2026/assets/papers/Paper_121/paper.pdf}
}
@ARTICLE{11400555,
author={Sarowar, Md Selim and Alnaasan, Manar and Kim, Sungho},
journal={IEEE Access},
title={C3G-VM6D: Data-Efficient C3G Vision Model Aided 6D Pose Estimation Based on RGB-D Data},
year={2026},
volume={14},
number={},
pages={34223-34238},
keywords={Pose estimation;Transformers;Point cloud compression;Three-dimensional displays;Probabilistic logic;Training;Visualization;Uncertainty;Foundation models;Feature extraction;Vision foundation model;cross-modal fusion;compact 3D Gaussian (C3G);vision transformer (ViT);6D object pose estimation;3D point cloud},
doi={10.1109/ACCESS.2026.3666439}}
* Corresponding authorAll publications