Devsinc is seeking a Senior AI/ML Engineer with over 4 years of experience, particularly in developing large language model (LLM) or generative AI applications. The ideal candidate will have a strong foundation in machine learning, along with expertise in Python backend development, APIs, retrieval-augmented generation (RAG), semantic search, AI agents, and scalable machine learning infrastructure. Experience working with both commercial and open-source AI models is essential. This role involves full ownership of the AI lifecycle—from experimentation and evaluation to deployment, optimization, and ongoing monitoring—ensuring solutions are reliable, scalable, secure, and cost-effective. The engineer will also contribute to architectural decisions, mentor junior team members, and collaborate with clients and cross-functional teams to drive measurable business outcomes.
Key Responsibilities
- Design, develop, and deploy AI/ML and LLM-based applications, including AI agents, tool-using systems, and human-in-the-loop workflows to address complex business challenges.
- Build scalable pipelines for training, fine-tuning, evaluation, and inference, incorporating experiment tracking and model versioning.
- Develop and optimize retrieval-augmented generation (RAG) and semantic search systems using embeddings, document chunking, vector search, reranking, and grounding techniques.
- Create backend APIs, microservices, and real-time inference services using Python frameworks such as FastAPI, Flask, or Django.
- Enhance model performance by improving quality, latency, throughput, and cost-efficiency through experimentation, hyperparameter tuning, quantization, batching, and caching.
- Implement MLOps best practices including automated testing, continuous integration/continuous deployment (CI/CD), and production monitoring supported by evaluation datasets, quality criteria, regression tests, and human review.
- Lead architectural decisions and cloud deployments ensuring scalability, reliability, security, and resource efficiency, while mitigating risks like prompt injection, data leakage, unauthorized access, and unsafe outputs.
- Evaluate emerging AI technologies and measure feature effectiveness using product analytics or A/B testing.
- Mentor engineers, collaborate effectively with technical and non-technical stakeholders, and document designs, experiments, and results comprehensively.
Required Qualifications
- Bachelor’s degree in Computer Science, Software Engineering, Data Science, or a related field.
- Minimum of 4 years of professional experience in AI/ML engineering, with proven ownership of production AI systems and hands-on development of LLM or generative AI applications.
- Strong production-level Python skills, with practical experience in PyTorch and/or TensorFlow, and solid understanding of machine learning concepts, neural networks, natural language processing (NLP), feature engineering, and model optimization.
- Experience integrating commercial or open-source LLMs, including prompt engineering, structured outputs, tool calling, context management, and managing model limitations.
- Hands-on expertise in building RAG or semantic search systems using embeddings, chunking, retrieval, reranking, and grounding, with vector search technologies such as pgvector, Pinecone, Weaviate, Qdrant, Milvus, or Elasticsearch.
- Proficiency in developing and deploying APIs, microservices, or inference services using FastAPI, Flask, Django, or similar frameworks, along with SQL and relational databases like PostgreSQL or MySQL.
- Experience deploying AI solutions on cloud platforms such as AWS, Azure, or Google Cloud, with familiarity in Git, Docker, automated testing, CI/CD pipelines, and MLOps practices including experiment tracking, model versioning, and monitoring.
- Understanding of AI evaluation methods, regression testing, human-in-the-loop review processes, and security and privacy considerations.
- Strong leadership skills with the ability to make technical decisions, guide engineering teams, and communicate clearly with clients and cross-functional stakeholders.
Preferred Qualifications
- Experience with AI frameworks such as LangChain, LlamaIndex, or LangGraph, and evaluation tools like LangSmith, Langfuse, or Ragas.
- Familiarity with Hugging Face Transformers, vLLM, Ollama, LoRA/PEFT fine-tuning, or self-hosted AI models.
- Knowledge of MLflow, Kubeflow, Kubernetes, Terraform, distributed systems, or GPU acceleration.
- Experience with data orchestration, asynchronous processing, caching, or messaging systems using Airflow, Redis, Celery, Kafka, or RabbitMQ.
- Exposure to advanced retrieval techniques, knowledge graphs, recommendation systems, computer vision, multimodal AI, or A/B testing and product analytics.
This role offers the opportunity to work on cutting-edge AI technologies within a dynamic environment, contributing to impactful projects that leverage the latest advancements in machine learning and generative AI.