This page grew out of BlueDot's Technical AI Safety course. I found the exercise useful, so I am sharing it here to document my commitment and contributions to this field.
Why AI Alignment?
I am focusing on technical AI alignment, particularly mechanistic interpretability, representation engineering, and methods for reliably understanding and controlling learned model behavior.
AI alignment matters to me because I believe our relationship with AI is about to change fundamentally. Today, we mostly treat AI as a tool. But as it becomes better than us at research, planning, education, and decision-making, delegating more responsibility to it will often be the most efficient choice.
Because of that, I think it is unacceptable that we could build and train these systems ourselves, become deeply dependent on them, and still not understand how they make important decisions or how their behavior might shift in unfamiliar situations. Alignment is therefore not only about preventing catastrophic failure. It is about ensuring that increasingly capable AI systems preserve core human values such as democracy, dignity, creativity, happiness, and independence.
Personally, I am drawn to this field because it offers relatively fast feedback loops. I can form hypotheses about model behavior, intervene directly, observe the effects, and iteratively refine my understanding.
I am also interested in how safety spans the entire AI lifecycle, from training-data curation and early detection of emerging capabilities to post-training alignment, mechanistic interpretability, deployment-time control, and continuous monitoring.
Target Roles
I am primarily interested in research scientist and research engineer roles in technical AI safety, particularly mechanistic interpretability, representation geometry, model steering and control, and empirical alignment. I am also exploring technical safety fellowships and PhD opportunities in mechanistic interpretability and AI safety, with a current goal of applying to US PhD programs for Fall 2027.
Experience
- Mechanistic interpretability @ UKP Lab, TU Darmstadt. I introduced Feature-Effect Geometry Analysis (FEGA), an unsupervised causal framework for studying how sparse autoencoder features affect model outputs across contexts. Our results show that clean one-dimensional downstream effects are rare: interpretable and causally relevant SAE features do not necessarily provide reliable steering directions. The work is currently under review at JMLR. Project · Paper · Code
- Efficient LLM inference @ M.Sc. thesis, MBZUAI. I developed and evaluated a training-free method for reusing internal representations across repeated token spans in Qwen2.5-3B. Shallow-layer representation reuse preserved model outputs while avoiding redundant computation.
- NLP research and dataset development. I led ViHOS, a dataset of 11,056 Vietnamese comments with span-level hate and offensive-language annotations and benchmark experiments. The work was published as a first-author paper at EACL 2023. Paper · Code
- ML systems and engineering. I worked as an AI Engineer at VinBigData on multilingual fine-grained NER, improving English and Bangla baselines for SemEval-2023 MultiCoNER II. At Fujairah Research Center, I built an internal RAG assistant spanning retrieval, prompt orchestration, backend logic, and response-generation workflows.
- Research background. I hold an M.Sc. in NLP from MBZUAI and a B.Sc. in Data Science from VNU-HCM University of Information Technology. My research spans mechanistic interpretability, sparse autoencoders, causal interventions, representation geometry, model steering, efficient generation, and NLP.
Engagement with AI Safety
- I completed the BlueDot Technical AI Safety course.
- I received an Honorable Mention in BlueDot Impact’s Technical AI Safety Puzzle #1 for an experiment on training a model to encode a semantic feature along a chosen nonlinear manifold using only three reserved channels. LessWrong write-up
- I am developing Feature-Effect Geometry Analysis (FEGA), an unsupervised causal framework for studying how SAE features affect model outputs across contexts. We find that interpretable SAE features rarely produce consistent one-dimensional effects, limiting their reliability as steering directions. Project · Paper
Actions
I will try to update this section once I have more information.