Artifical Intelligence (AI)Bengaluru, India
Head - AI Research
Location
Bengaluru, India
Compensation Tier
Leadership
Category
Artifical Intelligence (AI)
Company Overview
Our client is a Research lead organisation focusing on cutting-edge software, AI, and hardware innovation.
Position Overview
- Design, implement, and evaluate novel research prototypes for on-device personal LLM agents, with emphasis on adaptive test-time scaling, privacy-preserving memory consolidation, and RL-based personalization policies.
- Conduct literature surveys, identify research gaps, and formulate hypotheses and experimental plans aligned to the PersonalLLM research agenda.
- Build and maintain research codebases for training, fine-tuning, compression, and on-device inference of LLMs; reproduce baselines and benchmark against them.
- Design and run rigorous experiments (ablations, statistical significance, privacy-leakage measurement) and document findings in internal technical reports and publications.
- Prototype on-device inference pipelines (quantization, KV-cache management, speculative/adaptive decoding, early-exit, hardware-aware scheduling) and profile latency, memory, and energy on edge hardware.
- Develop privacy-preserving user-state representations (preference embeddings, user latent vectors, skill representations) and memory-consolidation mechanisms that discard raw interactions.
- Implement reinforcement-learning controllers for personalization policies (when to retrieve, reason longer, self-reflect, or update user representations) under cost/energy/privacy budgets.
- Collaborate with cross-functional engineering, product, and hardware teams to transition research prototypes into production-grade features.
Responsibilities
- Define the multi-year research roadmap for personal LLMs, aligning it with Samsung's device-intelligence vision and the evolving PersonalLLM agenda.
- Own end-to-end delivery of research objectives from problem framing and hypothesis design through prototyping, on-device validation, and production hand-off.
- Bridge the AS-IS → TO-BE transition: move the organization beyond distillation/pruning/quantization, RAG, and LoRA fine-tuning toward dynamic compression, semantic memory compression, neural prompt compression, learned KV-cache eviction, adaptive decoding, end-to-end agentic models with self-reflection/self-verification, private continual learning, and energy-aware inference planning.
- Establish research best practices: reproducibility, evaluation harnesses, privacy benchmarks, and on-device profiling standards.
- Represent SRIB in external research communities, open-source collaborations, and academic partnerships; build a pipeline of talent through mentoring and university engagement.
- Contribute to IP strategy by identifying patentable inventions and prior art.
Skills & Experience
- PhD/ Masters in a relevant AI field (e.g., Computer Science, Machine Learning, Artificial Intelligence, Natural Language Processing, or a closely related discipline).
- Experience
- Minimum 10 years of experience in a relevant research area (LLMs, on-device/edge AI, reinforcement learning, continual learning, privacy-preserving ML, or federated learning).
- Technical Skills
- Languages & Scripting: Python (expert), C/C++ (intermediate+), shell scripting; familiarity with on-device/mobile development (Java/Kotlin or Swift) is a plus.
- Deep-Learning & LLM Frameworks: PyTorch (expert), JAX (intermediate); Hugging Face Transformers, PEFT/LoRA, DeepSpeed, vLLM, TGI, TensorRT-LLM, ONNX Runtime, ExecuTorch, MLC-LLM, llama.cpp, MNN/TFLite.
- Model Compression & Efficient Inference: Distillation, pruning, quantization (INT8/INT4, weight-only, activation-aware), LoRA/QLoRA; dynamic compression, adaptive precision, task-specific extraction. KV-cache optimization (quantization, offloading, paging, learned eviction, semantic retention); FlashAttention, speculative decoding, early exit, adaptive decoding, hardware-aware scheduling.
- Context, Prompt & Memory: RAG, summarization, sliding-window context; semantic memory compression, neural memory tokens, neural prompt compression, latent intent representation.
- Reinforcement Learning & Test-Time Scaling: RLHF/RLAIF, PPO/DPO, policy-gradient methods, reward modeling; test-time compute scaling, adaptive inference controllers, cost-aware reasoning frameworks, personalized inference schedulers.
- Personalization, Continual & Privacy-Preserving Learning: Continual/lifelong learning, catastrophic-forgetting mitigation, private continual learning, lifelong memory; differential privacy, federated learning, secure aggregation.
Apply for this Role
You must be signed in to apply for this position.
Sign In to Apply