Hi! I am Reshmi Ghosh, a Senior Applied Scientist at Microsoft Turing in the Copilot, Agents and Core org (Previously at MSAI and Microsoft AI Acceleration Program under Office of the CTO). I currently serve as a Technical Lead to drive teams to build the next generation of productivity tools that are trustworthy and efficient. I am working in the areas of long-running asynchronous agentic tasks, computer-use coding harness, safety, personalization/memory, alignment, reasoning, and trustworthy-evaluations.
Most recently I was involved in R&D and productionization of several natural language and multimodal solutions for securing Copilot reasoning and enterprise RAG from direct and indirect prompt injections and worked on several cross-organization projects related to AI security. I have also contributed towards the post-training of AI-systems for aligning with human-preferences while ensuring reliable search experiences.
I have a rich experience of delivering high-impact features/products for Fortune 500 companies and end users. I also love prototyping and working on incubation projects, a desire stemming from my past experience of finishing a forward-looking Ph.D. from Carnegie Mellon University in the intersection of deep learning, natural language processing, and renewable energy policies.
For invited talks, conference presentations, and reviewing activities, see Professional Service.
I am looking forward to continue working in fast-paced teams and across multiple areas and disciplines/roles and growing to become an empathetic product leader. Email me at gsh.reshmi@gmail.com to discuss new opportunities. I have an employment-based green card and do not require visa sponsorship.
Education Background: I hold a Ph.D. from Carnegie Mellon University. I have always been passionate about applying machine learning, deep learning, economics to solving socio-technical problems, such as in the intersection of climate change and renewable energy integration.
With over 10+ years of innovating, leading, and strategizing AI solutions for nuanced user problems, I hope to continue building products and tools that can benefit the scientific community, and users equally.
Products/Software Highlights
A compact release timeline of what shipped, what hardened in production, and what became the next product platform bet.
2026 Memory and personalization become product primitives
The focus shifted from isolated features to durable user context, asynchronous work, and trustworthy long-running agents.
Platform
Memory
Personalization
- Built: Long-term memory systems that let Copilot retain user context across sessions while preserving privacy and trust boundaries.
- Shaped: Personalization signals that adapt agent behavior to individual users without compromising safety guardrails.
- Developed: Frontier evaluations for the long-running task harness and improved the reliability of asynchronous Copilot tasks.
- Released: Contributed to the release of computer-use agentic workflows.
- Explored: Memory-grounded reasoning patterns to reduce hallucinations and improve factual consistency in personalized workflows.
2025 The agentic release wave
This was the year product work became deeply agentic: orchestration, safety, grounding, and user trust all had to hold together at runtime.
Agents
Alignment
Runtime Safety
- Hardened: Fact alignment for a global audience while reducing over-refusal in product behavior.
- Released: Contributed to the safe launch of agentic workflows in M365 Copilots.
- Extended: Helped move Copilot from single-turn assistance toward more capable, multi-step agent experiences.
2023–2024 Responsible AI and first Copilot releases
This era connected evaluation research directly to product launches, from the first Copilot releases to enterprise-grade defenses for prompt injection and security failures.
Copilot
Evaluation
Responsible AI
- Led: State-of-the-art Responsible AI and trustworthy evaluation work for LLM and multimodal applications.
- Shipped: Early M365 Copilot capabilities, including Business Chat, as generative AI entered Microsoft productivity products.
- Translated: Research into production safeguards that later supported Azure Prompt Shields and related enterprise defenses.
2024 Securing AI search and productivity
The center of gravity was AI security: prompt injections, post-training defenses, and enterprise-grade trust for search and productivity systems.
Security
Post-Training
Enterprise AI
- Developed: Novel post-training methods for AI systems designed to prevent security breaches.
- Shipped: Azure Prompt Shields for third-party usage.
- Safeguarded: Automated search workflows against prompt injections and related security failures.
2021–2023 Office of the CTO / AI Incubation
Before the current Copilot era, the work centered on incubation: shipping foundational ML systems, experimenting aggressively, and learning how research ideas survive production constraints.
Incubation
ML Systems
Product Foundations
- Selected: Joined Microsoft’s AI incubation program at NERD with an approximately 0.1% acceptance rate.
- Highest impact: Contributed to the development and release of the first version of Copilot, covered here.
- Built: Anomaly detection systems for Azure core-service SLAs and intelligent commanding features for Microsoft Office.
- Developed: Probabilistic models behind Viva Topics and an LLM application plus evaluation framework for rich-text generation in a client application.
Research Highlights
Selected Papers
Selected list of published papers. For the full list, visit Google Scholar. Citation metrics as of August 2026: 477 citations, h-index 10, i10-index 10.
- [Mechanistic Interpretability] From RAGs to Rich Parameters: Probing How Language Models Utilize External Knowledge over Parametric Information for Factual Queries WSDM Companion, 2026
- [Alignment] Towards Bidirectional Human-AI Alignment: A Systematic Review for Clarifications, Framework, and Future Directions NeurIPS, 2025
- [Alignment] Bidirectional Human-AI Alignment: Emerging Challenges and Opportunities CHI, 2025
- [TrustworthyML/Safety] Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers EACL, 2025
- [Reasoning] Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysis Foundations of Reasoning Models Workshop at NeurIPS, 2025
- [Alignment] ValueCompass: A Framework of Fundamental Values for Human-AI Alignment Widening NLP Workshop at EMNLP, 2025
- [TrustworthyML/Safety] Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition NeurIPS, 2024; Best Paper
- [Mechanistic Interpretability] Quantifying reliance on external information over parametric knowledge during Retrieval Augmented Generation (RAG) using mechanistic analysis BlackBox NLP at EMNLP, 2024
- [Reasoning] Frontiers of Large Language Model-Based Agentic Systems: Construction, Efficacy and Safety CIKM Tutorial, 2024
- [Mechanistic Interpretability] On Surgical Fine-tuning for Language Encoder EMNLP, 2023
- [Reasoning] Topic Segmentation in the Wild: Towards Segmentation of Semi-structured & Unstructured Chats NeurIPS, 2022
- [Applied ML] Reconstruction of Long-Term Historical Electricity Demand Data ICML, 2021