Job Description:
On-site — Douglas Research Centre, McGill University (Montréal, QC)
Position Overview:
A central focus of this role is the development of modern representation-learning pipelines that support fine-grained feature extraction (e.g., micro-actions and subtle posture or kinematic signatures), multi-scale behaviour modelling (from individual frames to seconds, minutes, and higher-order behavioural motifs), and searchable video querying, enabling retrieval of “behaviour like this” through learned embeddings and large-scale indexing.
This is a joint position co-supervised by Drs. Majid Mohajerani and Pouya Bashivan at McGill University, bringing together large-scale behavioural video data with deep learning and computational neuroscience approaches.
The Mohajerani Lab, Neurodynamics and Behaviour Engineering Lab, is a multidisciplinary research group focused on behavioural neuroscience, large-scale sensing, and scalable machine-learning systems for behavioural neuroscience, with a focus on large-scale video understanding, multimodal representation learning, and self-supervised discovery of behaviour from long-duration (multi-month) recordings across multiple species.
The Bashivan Lab (Department of Physiology) develops computational models and neural network algorithms to explain, predict, and ultimately modulate brain responses and behaviour, drawing on machine learning, computational neuroscience, and cognitive science. The lab’s approach centers on building models that capture how neural populations support memory and behaviour, validated against electrophysiological, neuroimaging, and behavioural data.
Both labs are affiliated with Mila, the Quebec Artificial Intelligence Institute, providing access to a world-class AI research ecosystem. The successful applicant will be co-supervised by both PIs and will have the opportunity to collaborate with other Mila-affiliated researchers in computer vision, representation learning, foundation models, and large-scale optimization, while leading independent research at the frontier of AI for behavioural neuroscience.
Key Responsibilities:
- Design, train, and fine-tune state-of-the-art video representation models, including transformer-based architectures (e.g., ViT-style backbones, video transformers, masked video modelling, contrastive/self-supervised learning).
- Build behaviour embeddings that support semantic search and retrieval (query-by-example, text/attribute-guided retrieval if applicable, nearest-neighbour search in embedding space).
- Develop multi-scale modelling approaches: frame-level features → clip-level representations → longer-horizon structure (temporal pooling, hierarchical models, sequence models, event segmentation).
- Integrate pose, detection, segmentation, and tracking outputs into representation learning (late fusion / cross-attention / token-level conditioning).
- Implement and optimize large-scale data pipelines for training and inference (efficient data loading, caching, sharding, dataset versioning, and reproducibility).
- Establish rigorous evaluation protocols for behaviour modelling and retrieval (including representation quality, robustness, domain shift, and generalization across cohorts/setups).
- Collaborate with neuroscientists and engineers to translate models into usable tools for behavioural phenotyping and experimental workflows.
- Publish findings in peer-reviewed venues and present at conferences.
- Mentor trainees in modern ML/CV methods and reproducible research practices.
Preferred Technical Focus Areas (any strong subset is great):
- Video understanding: action recognition, temporal localization, long-range video modelling
- Self-supervised learning for video (masked prediction, contrastive learning, distillation)
- Representation learning + retrieval: embedding learning, metric learning, approximate nearest neighbour (ANN) indexing (FAISS / ScaNN), scalable search
- Multi-modal learning: combining video with sensor streams/metadata/pose tokens
- Behaviour segmentation/clustering: unsupervised discovery of motifs, hierarchical clustering, HMM/HSMM/AR-HMM style approaches (optional but valued)
- Efficient training/inference: mixed precision, distributed training, model acceleration/quantization (if relevant)
Qualifications:
- PhD in Computer Science, Electrical/Computer Engineering, Computational Neuroscience, or a related field.
- Strong track record in computer vision / deep learning, ideally with publications or substantial open-source work.
- Hands-on experience with transformer-based architectures and/or modern representation learning methods.
- Proficiency in Python and PyTorch (strongly preferred); familiarity with OpenCV, NumPy/Pandas, and practical ML tooling.
- Demonstrated ability to work with large-scale datasets (storage, throughput, data integrity, experiment tracking).
- Strong analytical thinking and debugging skills; comfortable working in a multidisciplinary setting.
- Prior experience in neuroscience/behavioral analysis is a plus but not required if your CV work aligns strongly with the modelling goals.
Salary and Benefits: Compensation:
$70,000–$80,000, commensurate with experience and qualifications, and subject to institutional approval. Initial 1-year appointment, with the possibility of extension for two more years based on performance and project needs.
How to Apply:
Please visit https://docs.google.com/forms/d/e/1FAIpQLSdgZkY2ltb1dix8ddKZD_s18lW86ushwTdxFnrQkmAFd1whdg/viewform and submit:
- A cover letter describing relevant experience and research interests (especially in video models, representation learning, and large-scale pipelines)
- A CV including publications, projects, and technical skills
- Contact information for 2–3 references
We are committed to fostering an equitable and inclusive environment and encourage applications from all qualified candidates, including women and members of visible minority groups.
