Ir al contenido principal
Senior LLM Engineer — enterprise RAG Todas las ofertas
Rol IA · Oferta abierta
Cohere
Tiempo completo Senior LLM platform Hybrid 88 Queens Quay W, Toronto, ON M5J 0B6, Canada · Hybrid Toronto, Canada CAD 150k–200k/year + equity Este anuncio está pensado para talento nativo en IA : habilidades y herramientas claras para saber si encajas antes de aplicar — y para reducir descartes por desajuste.
4 habilidades 2 herramientas
Compartir este rol
Copiar enlace
Modalidad Hybrid
Ubicación 88 Queens Quay W, Toronto, ON M5J 0B6, Canada · Hybrid
Compensación CAD 150k–200k/year + equity Nivel de experiencia SeniorBeneficios GPU access, equity, Toronto waterfront office. Cómo aplicar RAG or embedding pipeline with eval metrics and cost notes.
Publicada 31 jul 2026
Actualizada 3 ago 2026 Cohere is hiring a Senior Senior LLM Engineer — enterprise RAG on the LLM platform team (Hybrid · Toronto, Canada). This is a hands-on role in LLM & generative AI —not a generic “AI enthusiast” post.
You will own LLM features, retrieval, agents, and evaluation end to end: problem framing, offline evaluation, safe rollout, and iteration from production signals. Success is measured with clear metrics (quality, latency, cost, or business KPIs)—not slide decks alone.
Cohere builds enterprise-grade language models and APIs used globally from our Toronto headquarters.
Join as a Senior LLM Engineer to ship RAG, embeddings, and fine-tuning workflows for regulated customers.
What you will do
Ship production LLM and retrieval systems with offline evals, monitoring, and safe rollout. Work with product and domain experts on requirements and trade-offs. Improve data pipelines, labeling, and metrics tied to business KPIs. Document runbooks for on-call and customer-facing teams. What we look for
Strong Python and modern ML/LLM stack with shipped features. SQL, experimentation, and clear communication to stakeholders. Eligible to work in the listed country; professional English required. Python
Production Python for data pipelines, model code, APIs, and automated tests—not notebook-only workflows.
LangChain
Orchestrating LLM chains, tools, and retrieval with observable steps and failure handling.
OpenAI API
Calling frontier APIs with retries, cost controls, prompt versioning, and safety filters.
PyTorch
Training and inference with PyTorch: reproducible experiments, checkpointing, and deployment-minded exports.
Vector DBs Tool
Embedding storage, hybrid search, and freshness strategies for RAG and semantic retrieval.
Docker Tool
Containerized services and reproducible dev environments for ML and LLM workloads.
✓ You can show evidence in your application: RAG or embedding pipeline with eval metrics and cost notes. ✓ Several of these show up in recent shipped work—not only on your CV: Python, LangChain, OpenAI API, PyTorch, and Vector DBs. ✓ You have owned LLM features, retrieval, agents, and evaluation in production: debugging live issues, running postmortems, and iterating from real user or business signals. ✓ You are comfortable with the hybrid rhythm (Hybrid · Toronto, Canada)—onsite collaboration when it matters, deep work when remote. ✓ You explain trade-offs (quality, cost, latency, safety) in plain language—with numbers or examples, not buzzwords. ✓ You read the full listing (comp: CAD 150k–200k/year + equity, benefits in At a glance) and your expectations align. Lo que pide la empresa
RAG or embedding pipeline with eval metrics and cost notes.
Inicia sesión con cuenta candidato para aplicar
Regístrate, añade titular y bio; luego podrás adjuntar enlace o PDF de CV en Mi cuenta antes de aplicar. Así las candidaturas quedan ligadas a tu perfil.
Sugerencias
Talento IA que podría interesarte Perfiles ordenados por solape con habilidades y herramientas de este rol — útil si contratas equipo o comparas candidatos.
Ofertas similares Otros anuncios que comparten habilidades o herramientas — útil para stacks comparables o alternativas.
Ingénieur LLM — RAG enterprise LightOn
Indefinido (CDI) Tiempo completo Senior GenAI platform Hybrid 2 Rue de la Bourse, Paris, IDF 75002, France · Hybrid Paris, France EUR 70k–95k/year + BSPCE
Senior NLP Engineer — multilingual LLM fine-tuning Mediterra Language AI
Freelance / autónomo Tiempo completo Senior NLP Remote Spain · Remote Barcelona, Spain €65k–€82k + stock options Publicada el 19 abr 2026 Actualizada el 3 ago 2026
Mediterra — LoRA/QLoRA workflows for ES/EN/FR customer support bots; eval sets per locale.Remote Spain with quarterly Barcelona meetups.
Beneficios ·
Senior RAG Engineer — Spaces & inference Hugging Face
Indefinido (CDI) Tiempo completo Senior Inference & retrieval Hybrid 20 Rue de la Paix, Paris, IDF 75002, France · Hybrid Paris, France EUR 70k–95k/year + BSPCE
Agents Engineer — tool-use & planning H Company
Indefinido (CDI) Tiempo completo Senior Autonomous agents Hybrid 48 Rue de la Chaussée d'Antin, Paris, IDF 75009, France · Hybrid Paris, France EUR 80k–110k/year + BSPCE
Senior AI Developer — global news & trends WOP360
Tiempo completo Senior News intelligence Hybrid 101 Arch Street, Boston, MA 02110, United States · Hybrid Boston, United States $145k–$190k + equity Ver todas las ofertas →
Título
Senior LLM Engineer — enterprise RAG
Aplicar Pensado para equipos que contratan personas que trabajan con IA, no solo alrededor.
Los candidatos ven habilidades y herramientas al inicio; recibes candidaturas estructuradas en un solo panel.
Llega a profesionales de IA que buscan roles con expectativas basadas en pruebas.
Publicar oferta Publicada el 30 jun 2026 Actualizada el 3 ago 2026
LightOn déploie des LLM souverains et des systèmes RAG enterprise pour des grands comptes européens depuis Paris.Ingénieur LLM RAG (hybride Paris 2e) : retrieval hybride, fine-tuning, évaluation qualité et mise en production sur modèles LightOn.À propos de LightOnScale-up française spécialisée…
Beneficios · BSPCE, remote partiel, modèles souverains.
Python LangChain PyTorch OpenAI API
Herramientas Vector DBs Docker
Co-working stipend, conference budget.
Python PyTorch OpenAI API LangChain
Herramientas Docker Vector DBs
Publicada el 27 jul 2026
Actualizada el 3 ago 2026
Hugging Face héberge l'écosystème open-source ML le plus utilisé au monde — équipe produit à Paris.Senior RAG Engineer: retrieval, Spaces, et endpoints inference avec latence et qualité mesurées.MissionsAméliorer les pipelines RAG (chunking, embeddings, hybrid search).Publier des patterns…
Beneficios · BSPCE, remote partiel, communauté open-source.
Python PyTorch LangChain NLP
Herramientas Docker Kubernetes Vector DBs
Publicada el 25 jul 2026 Actualizada el 3 ago 2026
H Company construit des agents autonomes avec tool-use, planification et mémoire pour des workflows enterprise depuis Paris.Agents Engineer (hybride Paris 9e) : orchestrer LLM, outils internes et evals rigoureuses sur benchmarks internes et publics.À propos de H CompanyScale-up IA française…
Beneficios · BSPCE, benchmarks publics, équipe recherche appliquée.
Python LangChain OpenAI API PyTorch
Publicada el 3 jul 2026 Actualizada el 3 ago 2026
WOP360 is a global news platform covering 195 countries — politics, economy, technology, sport, and breaking briefings driven by real-world search trends.Senior AI Developer (Boston hybrid): build NLP and LLM systems that help editors detect trending topics, draft country-desk briefings, and…
Beneficios · Hybrid Boston, health coverage, conference budget, global newsroom access.
Python LangChain NLP OpenAI API SQL
Herramientas Vector DBs Docker React