RESEARCHER · BUILDER · EXPLORER

Hi, I’m
Shuvrajeet Das.Building agents
that learn to explore.

PhD student / Data Science & AI / IIT Madras

I study how intelligent agents learn through interaction. My work brings together reinforcement learning, generative models, and the engineering that turns an idea into an experiment.

COMPUTATION GRAPHCONCEPTUAL VIEW
Start with a question.

Make the assumptions explicit. Understand the idea before writing the implementation.

INTUITION → IMPLEMENTATION● INTERACTIVE

THE PERSON BEHIND THE EXPERIMENTS

Curiosity, with a working implementation.

I’m Shuvrajeet, a PhD student in Data Science and Artificial Intelligence at IIT Madras, exploring how intelligent agents can learn through interaction.

I enjoy taking an idea from equations to a working implementation—and investigating what happens when the experiments disagree with the intuition.

My research focus is deep reinforcement learning and diffusion-based exploration. The question that keeps me coming back: how can learned generative priors help agents explore more effectively?

Alongside research, I’m building self-play agents, game-playing systems, and continuous-control experiments. I’m also learning probabilistic machine learning, optimization, and GPU programming.

See what I’m building →

WHAT I THINK ABOUT

Six connected directions.

Learning. Decision-making. Exploration.

01 / LEARNING

Deep reinforcement learning

Actor–critic methods, off-policy learning, policy optimization, and sample efficiency.

02 / GENERATIVE PRIORS

Diffusion models for RL

Learning action priors and investigating their role in exploration.

03 / DISCOVERY

Exploration & uncertainty

Structured exploration, behavioral diversity, and probabilistic reasoning.

04 / COMPETITION

Self-play & planning

Game-playing agents, Monte Carlo tree search, and learning through competition.

05 / INTERACTION

Continuous control & robotics

Sequential decision-making in control tasks and embodied environments.

06 / SYSTEMS

Efficient learning systems

Vectorized rollouts, GPU computation, and reproducible training pipelines.

02 / QUESTIONS WORTH EXPLORING

At the research frontier.

FROM THE PROFILE README

RESEARCH SPOTLIGHT

Can a learned prior
lead to better exploration?

I’m investigating diffusion models as a source of structured exploration for reinforcement learning: can a model propose useful exploratory actions while an RL agent learns how to refine and use them?

I’m interested in how those proposals interact with a learned actor, and how to evaluate their contribution through verified baselines and controlled experiments.

Read the research direction ↗
Collected experienceLearned action priorStructured explorationOngoing research · benefits not yet established

WHAT I’M BUILDING

From agents to robots.

All my repositories ↗

Project descriptions come from my profile README. This website’s repository does not contain their implementations or experiment results.

INSIDE THE LABORATORY

Algorithm explorer.

Follow the implementation.
Understand the computation.

Reading repository inventory…

Loading…
What does this catalog include?

Research projects mentioned in my profile are shown separately above. They are not counted as local implementations.

03 / THE COMPUTE LAYER

One idea. Many operations.

Hardware is part of the question.

Understand what
the machine is doing.

The README describes TensorFlow pipelines, vectorized environments, and multi-GPU training as engineering interests.

This diagram is conceptual. Device support, memory use, and speed depend on the actual implementation and must be checked per experiment.

NO LOCAL HARDWARE BENCHMARKS
TensorFlow tensor operations
CPU

General-purpose compute

Control flow and numerical operations

GPU

Parallel compute

Supported operations across many elements

Conceptual execution paths · not a performance comparison

04 / EVIDENCE OVER ASSUMPTIONS

The experiment desk.

Experiment results · coming soon

No local results or plots were found in the inspected checkout. Research directions in the README are not measured outcomes.

AWAITING ARTIFACTS

05 / OPEN THE NOTEBOOK

A map of the work.

View repository ↗
TheUnsolvedDev / project inventory

READING PATH

Start at the source.

The README is the authoritative introduction to the research. Implementation-specific learning paths become available as source directories are added.

Open the README ↗

HOW I BUILD

Less hype.
More experiments.

Let the ablations speak.

I work with TensorFlow-first model development, custom training loops, and reusable research components. I care about understanding which components contribute, and accounting for stability and compute alongside reward.

01

Make the idea concrete

Take an idea from equations to a working implementation. Separate models, data pipelines, training, and evaluation.

02

Build for reproducibility

Control randomness, track experimental settings, and design TensorFlow code with graph execution in mind.

03

Question the result

Verify baselines before adding complexity. Use ablations to investigate what actually contributes.

ON MY WORKBENCH

Research meets engineering.

Tools and setup from my profile.

Machine learning & research

PythonTensorFlowscikit-learnOpenCV

Programming & systems

C / C++CUDABashLinuxGit

MLOps & data workflows

DockerGitHub ActionsMLflowDVCRaySparkAirflowKafkaMySQLMongoDB

Hardware & creative tools

ArduinoBlenderUnityRoboticsAnimation

MY SETUP · AS DOCUMENTED

Training agents. Debugging kernels.

And occasionally remembering the GPUs can run games too.

OS
Manjaro Linux
GPU
2 × RTX 4060 Ti
Memory
128 GB RAM
Storage
NVMe SSD

OPEN TO RESEARCH DISCUSSIONS & COLLABORATIONS

Working on a
similar question?

If you’re thinking about exploration, diffusion models, or learning agents, I’d love to exchange ideas.

shuvrajeet17@gmail.com