Projects

5
Bootcamps & Courses
4
Research Projects
4
Industry Projects
21
Skills

Bootcamps & Courses

Extensive hands-on training programs and graduate coursework

5 programs

Graduate Course · Fall 2026 · Instructor: Jiatao Gu

Graduate course on learned representations and predictors of environment dynamics for perception, planning, and decision making, and how they support control and reasoning across reinforcement learning, video and 3D, multimodal agents, and robotics. Twenty-three lectures spanning probabilistic formulations of world models, environments and rollouts, state-space models, representation learning, and generative foundations (latent-variable, adversarial, autoregressive, diffusion and flow matching, normalizing flows).

  • Sequential generative world models; model-based reinforcement learning, planning, and control
  • Video world models: generative architectures, rollout, interaction, memory, efficiency
  • Spatial world models: 3D geometry, 4D dynamics, neural physics and learned physical dynamics
  • World models for robot learning: simulation, sim-to-real, vision-language-action and world-action models
  • LLMs as world models: simulation, state tracking, grounding; reasoning with test-time computation; world models for digital agents
  • Evaluating world models: utility, robustness, and open problems

Course page

3-Month Extensive Bootcamp · Vizuara AI Lab

Completed vision‑language foundations, open‑source VLAs like Pi0, SmolVLA, and OpenVLA, Diffusion Policy for robotic manipulation, world‑model‑based simulation, SO‑101 integration, and edge deployment, culminating in a capstone project.

3-Month Extensive Bootcamp · Vizuara AI Lab

Gained hands-on experience with large language models, context foundations, instructional prompting, and building dynamic personal and voice agents. And so, Implemented Retrieval-Augmented Generation (WRITE & SELECT), tool use and MCP-based workflows, and advanced memory architectures for production-grade AI agents.

3-Month Extensive Bootcamp · Vizuara AI Lab

Gained hands-on experience in building the critical infrastructure required to transform LLMs into production-grade autonomous agents. Implemented custom agent loops (ReAct pattern), sandboxed tool execution systems with permission gates, memory/context engines, and control planes for error-resilient execution (checkpoints and replay).

3-Month Extensive Bootcamp · Vizuara AI Lab

Implemented and coded from scratch models like ViT, DeiT, Swin, DETR, Mask2Former, SAM, TimeSformer, VideoMAE, CLIP, BLIP, Flamingo, LLaVA, and Stable Diffusion. As well, projects on object detection and segmentation to multimodal chat systems and generative image pipelines.

Skills

Tools, frameworks & languages I work with

21 skills
Frameworks & Libraries
PyTorchHugging FaceTensorFlowMatplotlibOpen3DROS2CUDAOpenCV
Research Tools & Hardware
Blender (3D Modeling)Isaac SimIsaac LabLaTeXWorld ModelsVLAsJetson AGX OrinRGBD+LiDAR sensor
Languages
Bengali: NativeEnglish: FluentKorean: IntermediateHindi: FluentChinese (Mandarin): Elementary

Industry Research Collaboration

Funded projects with industry & research institutes

4 projects · 4 partners
Towards General-Purpose Physical AI: Infrastructure, World Models, Spatial Reasoning, and Robot Manipulation
  • PI: Prof. Sungho Kim
  • Building shared infrastructure and benchmarks for general-purpose physical AI, spanning perception, world-model learning, and spatial reasoning for robot manipulation
  • Integrating action-conditioned world models with spatial reasoning modules to enable robust, generalizable manipulation policies across diverse robotic platforms
Agentic Policy & World Models for Autonomous Mobile Robots
  • Building action-conditioned latent world models and agentic planning loops with latent lookahead to verify candidate actions in imagination before physical execution
  • Deploying vision-language-action policies and sim-to-real pipelines with Isaac Lab and ROS 2 on edge mobile robot platforms
Pose Estimation with VLM integration for RGB-D, LiDAR & RADAR Multi Modal fusion for workers safety and Industry automation
  • Developing multi-sensor fusion pipelines (RGB-D, LiDAR, and RADAR) to estimate 3D object poses and segment dynamic safety boundaries
  • Integrating vision-language models (VLMs) to semantically analyze industrial scenes and identify hazardous worker-robot proximity events
3D Human Pose Estimation with VLM integration for Healthcare
  • Implementing 3D human skeleton and mesh recovery models from monocular RGB-D video streams to track patient rehabilitation progress
  • Fusing spatial joint trajectories with vision-language models (VLMs) to generate clinical reports on mobility and gait analysis

Research & Applied

Research prototypes and applied AI systems

4 projects

Building Personal AI News Aggregator with Agentic AI (Claude Code)

Building Personal AI News Aggregator with Agentic AI (Claude Code)

  • Aggregates AI updates from curated high-quality sources
  • Filters noise
  • Adapts to personal preferences
  • Summarizes intelligently
  • Allows feedback (thumbs up / thumbs down style)
  • Improves over time using an agent-based backend

Repository

Personal Chatbot by instructions fine-tuning LLM with open-source dataset

Personal Chatbot by instructions fine-tuning LLM with open-source dataset

we run the LLM fine-tuning loop on the instruction dataset. We demonstrate how fine-tuning can improve LLM performance while following instructions.

Repository

Fine-tuning LLM for Classification task, email Spam or ham, with UC Irvine dataset

Fine-tuning LLM for Classification task, email Spam or ham, with UC Irvine dataset

we run the LLM fine-tuning loop on the email spam or not spam(ham) dataset. We demonstrate how fine-tuning can improve LLM performance in the classification tasks.

Repository

My Research Proposal for RISE Project: 2026 (2nd year)

My Research Proposal for RISE Project: 2026 (2nd year)

OpenVLA 모델을 파인튜닝하고, 추론 속도 향상, 스케일링 특성 개선, 제로샷 일반화 성능 및 장기 시퀀스(장시간 작업) 수행 능력을 강화하기 위한 고성능·경량화 아키텍처를 설계한다. 모델 기반 및 시뮬레이션 기반 접근을 결합한 SOTA 수준의 Efficient-VLA 네트워크를 구축하고, 벤치마크 평가를 수행하여 ACCV 2026(일본 오사카)에 논문을 제출한다. 또한, Sim-to-Real 및 Real-to-Sim 전이 학습 기법을 활용하여 5지(5-finger) 로봇 그리퍼와 NVIDIA Jetson AGX Orin 플랫폼에 비전-언어-행동(VLA) 모델을 적용한다. 다중 센서 융합을 통해 로봇 인지 및 조작 데이터를 대규모로 확장하고, 이를 기반으로 CVPR, ICCV, ECCV 등 최상위 국제학회 및 저널에 연구 성과를 발표한다.

Certifications

MOOC & workshop credentials

Copied to clipboard!