Projects
5
Bootcamps & Courses
4
Research Projects
4
Industry Projects
21
Skills
Bootcamps & Courses
Extensive hands-on training programs and graduate coursework
Graduate course on learned representations and predictors of environment dynamics for perception, planning, and decision making, and how they support control and reasoning across reinforcement learning, video and 3D, multimodal agents, and robotics. Twenty-three lectures spanning probabilistic formulations of world models, environments and rollouts, state-space models, representation learning, and generative foundations (latent-variable, adversarial, autoregressive, diffusion and flow matching, normalizing flows).
- Sequential generative world models; model-based reinforcement learning, planning, and control
- Video world models: generative architectures, rollout, interaction, memory, efficiency
- Spatial world models: 3D geometry, 4D dynamics, neural physics and learned physical dynamics
- World models for robot learning: simulation, sim-to-real, vision-language-action and world-action models
- LLMs as world models: simulation, state tracking, grounding; reasoning with test-time computation; world models for digital agents
- Evaluating world models: utility, robustness, and open problems
Completed vision‑language foundations, open‑source VLAs like Pi0, SmolVLA, and OpenVLA, Diffusion Policy for robotic manipulation, world‑model‑based simulation, SO‑101 integration, and edge deployment, culminating in a capstone project.
Gained hands-on experience with large language models, context foundations, instructional prompting, and building dynamic personal and voice agents. And so, Implemented Retrieval-Augmented Generation (WRITE & SELECT), tool use and MCP-based workflows, and advanced memory architectures for production-grade AI agents.
Gained hands-on experience in building the critical infrastructure required to transform LLMs into production-grade autonomous agents. Implemented custom agent loops (ReAct pattern), sandboxed tool execution systems with permission gates, memory/context engines, and control planes for error-resilient execution (checkpoints and replay).
Implemented and coded from scratch models like ViT, DeiT, Swin, DETR, Mask2Former, SAM, TimeSformer, VideoMAE, CLIP, BLIP, Flamingo, LLaVA, and Stable Diffusion. As well, projects on object detection and segmentation to multimodal chat systems and generative image pipelines.
Skills
Tools, frameworks & languages I work with
Frameworks & Libraries
PyTorchHugging FaceTensorFlowMatplotlibOpen3DROS2CUDAOpenCV
Research Tools & Hardware
Blender (3D Modeling)Isaac SimIsaac LabLaTeXWorld ModelsVLAsJetson AGX OrinRGBD+LiDAR sensor
Languages
Bengali: NativeEnglish: FluentKorean: IntermediateHindi: FluentChinese (Mandarin): Elementary
Industry Research Collaboration
Funded projects with industry & research institutes
Towards General-Purpose Physical AI: Infrastructure, World Models, Spatial Reasoning, and Robot Manipulation
- PI: Prof. Sungho Kim
- Building shared infrastructure and benchmarks for general-purpose physical AI, spanning perception, world-model learning, and spatial reasoning for robot manipulation
- Integrating action-conditioned world models with spatial reasoning modules to enable robust, generalizable manipulation policies across diverse robotic platforms
Agentic Policy & World Models for Autonomous Mobile Robots
- Building action-conditioned latent world models and agentic planning loops with latent lookahead to verify candidate actions in imagination before physical execution
- Deploying vision-language-action policies and sim-to-real pipelines with Isaac Lab and ROS 2 on edge mobile robot platforms
Pose Estimation with VLM integration for RGB-D, LiDAR & RADAR Multi Modal fusion for workers safety and Industry automation
- Developing multi-sensor fusion pipelines (RGB-D, LiDAR, and RADAR) to estimate 3D object poses and segment dynamic safety boundaries
- Integrating vision-language models (VLMs) to semantically analyze industrial scenes and identify hazardous worker-robot proximity events
3D Human Pose Estimation with VLM integration for Healthcare
- Implementing 3D human skeleton and mesh recovery models from monocular RGB-D video streams to track patient rehabilitation progress
- Fusing spatial joint trajectories with vision-language models (VLMs) to generate clinical reports on mobility and gait analysis
Research & Applied
Research prototypes and applied AI systems
- Aggregates AI updates from curated high-quality sources
- Filters noise
- Adapts to personal preferences
- Summarizes intelligently
- Allows feedback (thumbs up / thumbs down style)
- Improves over time using an agent-based backend
we run the LLM fine-tuning loop on the instruction dataset. We demonstrate how fine-tuning can improve LLM performance while following instructions.
we run the LLM fine-tuning loop on the email spam or not spam(ham) dataset. We demonstrate how fine-tuning can improve LLM performance in the classification tasks.
OpenVLA 모델을 파인튜닝하고, 추론 속도 향상, 스케일링 특성 개선, 제로샷 일반화 성능 및 장기 시퀀스(장시간 작업) 수행 능력을 강화하기 위한 고성능·경량화 아키텍처를 설계한다. 모델 기반 및 시뮬레이션 기반 접근을 결합한 SOTA 수준의 Efficient-VLA 네트워크를 구축하고, 벤치마크 평가를 수행하여 ACCV 2026(일본 오사카)에 논문을 제출한다. 또한, Sim-to-Real 및 Real-to-Sim 전이 학습 기법을 활용하여 5지(5-finger) 로봇 그리퍼와 NVIDIA Jetson AGX Orin 플랫폼에 비전-언어-행동(VLA) 모델을 적용한다. 다중 센서 융합을 통해 로봇 인지 및 조작 데이터를 대규모로 확장하고, 이를 기반으로 CVPR, ICCV, ECCV 등 최상위 국제학회 및 저널에 연구 성과를 발표한다.
Certifications
MOOC & workshop credentials