Skip to main content
Senior LLM Engineer — enterprise RAG All open roles
AI role · Open offer
Cohere
Full-time Senior LLM platform Hybrid 88 Queens Quay W, Toronto, ON M5J 0B6, Canada · Hybrid Toronto, Canada CAD 150k–200k/year + equity This listing is optimized for AI-native talent : clear skills and tools so candidates know if they fit before they apply—and so you spend less time screening mismatches.
Workplace Hybrid
Location 88 Queens Quay W, Toronto, ON M5J 0B6, Canada · Hybrid
Compensation CAD 150k–200k/year + equity Benefits & perks GPU access, equity, Toronto waterfront office. How to apply RAG or embedding pipeline with eval metrics and cost notes.
Listed Jul 31, 2026
Updated Aug 3, 2026 Cohere is hiring a Senior Senior LLM Engineer — enterprise RAG on the LLM platform team (Hybrid · Toronto, Canada). This is a hands-on role in LLM & generative AI —not a generic “AI enthusiast” post.
You will own LLM features, retrieval, agents, and evaluation end to end: problem framing, offline evaluation, safe rollout, and iteration from production signals. Success is measured with clear metrics (quality, latency, cost, or business KPIs)—not slide decks alone.
Cohere builds enterprise-grade language models and APIs used globally from our Toronto headquarters.
Join as a Senior LLM Engineer to ship RAG, embeddings, and fine-tuning workflows for regulated customers.
What you will do
Ship production LLM and retrieval systems with offline evals, monitoring, and safe rollout. Work with product and domain experts on requirements and trade-offs. Improve data pipelines, labeling, and metrics tied to business KPIs. Document runbooks for on-call and customer-facing teams. What we look for
Strong Python and modern ML/LLM stack with shipped features. SQL, experimentation, and clear communication to stakeholders. Eligible to work in the listed country; professional English required. Python
Production Python for data pipelines, model code, APIs, and automated tests—not notebook-only workflows.
LangChain
Orchestrating LLM chains, tools, and retrieval with observable steps and failure handling.
OpenAI API
Calling frontier APIs with retries, cost controls, prompt versioning, and safety filters.
PyTorch
Training and inference with PyTorch: reproducible experiments, checkpointing, and deployment-minded exports.
Vector DBs Tool
Embedding storage, hybrid search, and freshness strategies for RAG and semantic retrieval.
Docker Tool
Containerized services and reproducible dev environments for ML and LLM workloads.
✓ You can show evidence in your application: RAG or embedding pipeline with eval metrics and cost notes. ✓ Several of these show up in recent shipped work—not only on your CV: Python, LangChain, OpenAI API, PyTorch, and Vector DBs. ✓ You have owned LLM features, retrieval, agents, and evaluation in production: debugging live issues, running postmortems, and iterating from real user or business signals. ✓ You are comfortable with the hybrid rhythm (Hybrid · Toronto, Canada)—onsite collaboration when it matters, deep work when remote. ✓ You explain trade-offs (quality, cost, latency, safety) in plain language—with numbers or examples, not buzzwords. ✓ You read the full listing (comp: CAD 150k–200k/year + equity, benefits in At a glance) and your expectations align. What the employer asked for
RAG or embedding pipeline with eval metrics and cost notes.
Sign in with a candidate account to apply
Register, add a headline and bio, then you can attach a CV link or PDF on My account before you apply. That keeps applications tied to your profile and makes it easier for employers to review your background.
Suggested
AI talent you may also like Profiles ranked by overlap with this role's skills and tools—handy if you're hiring for a team or comparing backup candidates.
Similar open roles Other listings that share skills or tools with this one—useful if you want comparable stacks or backup options.
Ingénieur LLM — RAG enterprise LightOn
Permanent (CDI) Full-time Senior GenAI platform Hybrid 2 Rue de la Bourse, Paris, IDF 75002, France · Hybrid Paris, France EUR 70k–95k/year + BSPCE
Senior NLP Engineer — multilingual LLM fine-tuning Mediterra Language AI
Freelance Full-time Senior NLP Remote Spain · Remote Barcelona, Spain €65k–€82k + stock options Listed Apr 19, 2026 Updated Aug 3, 2026
Mediterra — LoRA/QLoRA workflows for ES/EN/FR customer support bots; eval sets per locale.Remote Spain with quarterly Barcelona meetups.
Perks ·
Senior RAG Engineer — Spaces & inference Hugging Face
Permanent (CDI) Full-time Senior Inference & retrieval Hybrid 20 Rue de la Paix, Paris, IDF 75002, France · Hybrid Paris, France EUR 70k–95k/year + BSPCE
Agents Engineer — tool-use & planning H Company
Permanent (CDI) Full-time Senior Autonomous agents Hybrid 48 Rue de la Chaussée d'Antin, Paris, IDF 75009, France · Hybrid Paris, France EUR 80k–110k/year + BSPCE
Senior AI Developer — global news & trends WOP360
Full-time Senior News intelligence Hybrid 101 Arch Street, Boston, MA 02110, United States · Hybrid Boston, United States $145k–$190k + equity Browse all open roles →
Title
Senior LLM Engineer — enterprise RAG
Apply Built for teams hiring people who work with AI, not only around it.
Candidates see skills and tools upfront; you get structured applications in one dashboard.
Reach AI practitioners browsing open roles with proof-first expectations.
Post a job Listed Jun 30, 2026 Updated Aug 3, 2026
LightOn déploie des LLM souverains et des systèmes RAG enterprise pour des grands comptes européens depuis Paris.Ingénieur LLM RAG (hybride Paris 2e) : retrieval hybride, fine-tuning, évaluation qualité et mise en production sur modèles LightOn.À propos de LightOnScale-up française spécialisée…
Perks · BSPCE, remote partiel, modèles souverains.
Python LangChain PyTorch OpenAI API
Co-working stipend, conference budget.
Python PyTorch OpenAI API LangChain
Listed Jul 27, 2026
Updated Aug 3, 2026
Hugging Face héberge l'écosystème open-source ML le plus utilisé au monde — équipe produit à Paris.Senior RAG Engineer: retrieval, Spaces, et endpoints inference avec latence et qualité mesurées.MissionsAméliorer les pipelines RAG (chunking, embeddings, hybrid search).Publier des patterns…
Perks · BSPCE, remote partiel, communauté open-source.
Python PyTorch LangChain NLP
Tools Docker Kubernetes Vector DBs
Listed Jul 25, 2026 Updated Aug 3, 2026
H Company construit des agents autonomes avec tool-use, planification et mémoire pour des workflows enterprise depuis Paris.Agents Engineer (hybride Paris 9e) : orchestrer LLM, outils internes et evals rigoureuses sur benchmarks internes et publics.À propos de H CompanyScale-up IA française…
Perks · BSPCE, benchmarks publics, équipe recherche appliquée.
Python LangChain OpenAI API PyTorch
Listed Jul 3, 2026 Updated Aug 3, 2026
WOP360 is a global news platform covering 195 countries — politics, economy, technology, sport, and breaking briefings driven by real-world search trends.Senior AI Developer (Boston hybrid): build NLP and LLM systems that help editors detect trending topics, draft country-desk briefings, and…
Perks · Hybrid Boston, health coverage, conference budget, global newsroom access.
Python LangChain NLP OpenAI API SQL
Tools Vector DBs Docker React