PhD Machine Learning & Vision Scientist
Senior Machine Learning & Vision Scientist
About AIM Robots
AIM Robots build AI-native industrial robots that learn from video, understand complex manufacturing environments, and execute skilled physical tasks autonomously.
We are seeking exceptional PhD-level researchers—especially from world-class groups such as Yann LeCun’s NYU/FAIRlab, Stanford U, OpenAI, GoogleDeepMind—to define how robots see, learn, plan, and act in real production environments.
This is a hybrid research–engineering role with both scientific freedom and immediate real-world deployment impact.
Role Type
Research–Engineering Hybrid
Full-time or Part-time
Founding Scientist Track
What You Will Lead
You will architect AIM Robots’ next-generation perception, representation learning, real world knowledge abstration and world-modeling foundation.
You will design and deploy:
- Multi-modal models
- Video understanding systems
- World models for prediction and planning
- Robust policy-learning pipelines for autonomous robot behavior in real factories
You will work closely with the founding team to shape both the research roadmap and the production AI stack.
Core FocusAreas
-
Perception & Video Understanding
- Design 2D/3D perception pipelines for complex, cluttered manufacturing environments
- Fuse information from multiple cameras, RGB-D sensors, and other modalities
- Develop video models that understand temporal structure and operator actions
- Convert raw human demonstrations into structured, machine-interpretable representations
-
Representation Learning & Self-Supervision
- Build scalable self-supervised learning (SSL) pipelines
- (e.g., JEPA / I-JEPA,YOLO, DINO, SAM, MAE, MoCo)
- Develop efficient representation backbones optimized for real-time performance and reliability
- Explore structured intermedidate and other advanced representation-learning and understanding approaches that tightly link perception and controls
- Build scalable self-supervised learning (SSL) pipelines
-
World Models & Autonomous Behavior
- Develop multi-modal world models that combine perception, dynamics, and control
- Build action-conditioned generative models (e.g., transformers, diffusion) for prediction and planning
- Enable multi-step reasoning and planning for complex physical tasks
- Drive generalization across workflows, products, and factories
-
Imitation Learning, RL & PolicyGeneration
- Develop behaviorcloning, DAgger, and offline RL pipelines from demonstration data
- Build diffusion-policy or related generative control models for fine manipulation and robust policies
- Extract reusable, generalizable behaviors from human demonstration video
- Collaborate with robotics engineers to deploy and iterate policies in production environments
-
Sim2Real, Synthetic Data & TrainingInfrastructure
- Design synthetic data generation workflows for perception and policy learning
- Use domain randomization and physics-based variation to improve robustness under distribution shifts
- Integrate digitaltwins and simulation environments to accelerate large-scale training
- Create tools and pipelines for training, evaluating, and monitoring multi-modal models at scale
Qualifications Required
- PhD (or equivalent research track) in Machine Learning, Computer Vision, Robotics, or a closely related field
- Expertise in self-supervised learning, representation learning, video modeling, or world models
- Strong fundamentals in 2D/3D perception and modern deep learning
- Experience training and scaling deep learning systems on real data
- Familiarity with robotics, embodied AI, or ML systems that operate in the physical world
Nice to Have
- Experience with Deep Learning models, VLM/VLA or world-model research
- Strong engineering skills in PyTorch, C++, CUDA, or Triton
- Background in top-tier research labs (e.g., NYU, FAIR, DeepMind,Google Brain)
- Experience working with large-scale multi-modal datasets or real robot data
Compensation & Track
We offer a competitive salary and meaningful equity. The Founding Scientist track includes:
- Enhanced equity grants
- Optional participation in revenue tied to developed AI modules and deployed systems
Why AIM Robotics
- Your research goes directly into real U.S. factories and production systems
- Access to large-scale multi-modal data: video, depth, sensor streams, and more
- Opportunity to design the intelligence stack for next-generation industrial robots
- Blend of academic-style research freedom with fast, pragmatic engineering execution
- High-impact role influencing technical strategy, system architecture, and hiring