Devsinc is seeking a Senior AI/ML Engineer with over 4 years of professional experience, particularly in developing large language models (LLM) or generative AI applications. The ideal candidate will have a solid foundation in machine learning principles and strong expertise in building Python backend services and APIs. Experience with retrieval-augmented generation (RAG), semantic search, AI agents, and scalable machine learning infrastructure using both commercial and open-source models is essential. This role involves taking full ownership of the AI lifecycle—from experimentation and evaluation to deployment, optimization, and ongoing monitoring—ensuring solutions are reliable, scalable, secure, and cost-effective. The engineer will also contribute to architectural decisions, mentor junior engineers, and collaborate closely with clients and cross-functional teams to drive measurable business outcomes.
Key Responsibilities
- Design, develop, and deploy AI/ML and LLM-based applications, including AI agents, tool-using systems, and human-in-the-loop workflows aimed at solving complex business challenges.
- Build scalable pipelines for training, fine-tuning, evaluation, and inference, incorporating experiment tracking and model versioning.
- Develop and optimize RAG and semantic search systems using techniques such as embeddings, document chunking, vector search, reranking, and grounding.
- Create backend APIs, microservices, and real-time inference services using Python frameworks like FastAPI, Flask, or Django.
- Enhance model performance by improving quality, latency, throughput, and cost-efficiency through experimentation, hyperparameter tuning, quantization, batching, and caching.
- Implement MLOps best practices including automated testing, continuous integration/continuous deployment (CI/CD), and production monitoring supported by evaluation datasets, regression tests, and human review.
- Lead architectural decisions and cloud deployment strategies to ensure scalability, reliability, security, and resource efficiency, while mitigating risks such as prompt injection, data leakage, unauthorized access, and unsafe outputs.
- Evaluate emerging AI technologies and assess feature effectiveness using product analytics or A/B testing.
- Mentor engineering teams, collaborate with both technical and non-technical stakeholders, and maintain thorough documentation of designs, experiments, and results.
Required Qualifications
- Bachelor’s degree in Computer Science, Software Engineering, Data Science, or a related discipline.
- Minimum of 4 years of professional experience in AI/ML engineering with proven ownership of production AI systems and hands-on development of LLM or generative AI applications.
- Strong production-level Python programming skills, with practical experience in PyTorch and/or TensorFlow. Solid understanding of machine learning, neural networks, natural language processing (NLP), feature engineering, and model optimization.
- Experience integrating commercial or open-source LLMs, including prompt engineering, structured outputs, tool calling, context management, and managing model limitations.
- Hands-on experience building RAG or semantic search systems using embeddings, chunking, retrieval, reranking, and grounding, with vector search technologies such as pgvector, Pinecone, Weaviate, Qdrant, Milvus, or Elasticsearch.
- Proficiency in developing and deploying APIs, microservices, or inference services using FastAPI, Flask, Django, or equivalent frameworks, along with SQL and relational databases like PostgreSQL or MySQL.
- Experience deploying AI solutions on cloud platforms such as AWS, Azure, or Google Cloud. Familiarity with Git, Docker, automated testing, CI/CD pipelines, and MLOps tools including experiment tracking, model versioning, and monitoring.
- Strong understanding of AI evaluation methodologies, regression testing, human-in-the-loop review processes, and awareness of security and privacy risks.
- Ability to make technical decisions independently, guide engineering teams, and communicate effectively with clients and cross-functional stakeholders.
Preferred Qualifications and Additional Skills
- Experience with AI frameworks such as LangChain, LlamaIndex, or LangGraph, and evaluation platforms like LangSmith, Langfuse, or Ragas.
- Familiarity with Hugging Face Transformers, vLLM, Ollama, LoRA/PEFT fine-tuning, or managing self-hosted models.
- Knowledge of MLflow, Kubeflow, Kubernetes, Terraform, distributed systems, or GPU acceleration technologies.
- Experience with data orchestration, asynchronous processing, caching, or messaging systems using tools such as Airflow, Redis, Celery, Kafka, or RabbitMQ.
- Exposure to advanced retrieval techniques, knowledge graphs, recommendation systems, computer vision, multimodal AI, or product analytics including A/B testing.
This opportunity offers the chance to work on cutting-edge AI technologies in a dynamic environment, contributing to impactful projects while growing technical leadership skills.