We are seeking a skilled software engineer to develop and deploy generative AI applications, including large language model (LLM)-powered features, retrieval-augmented generation (RAG) pipelines, and enterprise search solutions. This role focuses on building production-ready AI systems that go beyond prototypes, delivering real-world impact through innovative AI-driven products. The ideal candidate will work closely with cross-functional teams to translate client requirements into scalable GenAI features while optimizing performance, cost, and user experience.
Key Responsibilities
- Develop generative AI applications such as LLM-powered features, copilot/chat experiences, and enterprise search tools.
- Design and implement RAG pipelines, including chunking strategies, embedding selection, hybrid retrieval, re-ranking, and GraphRAG for structured data retrieval.
- Fine-tune and adapt models using techniques like LoRA and QLoRA when prompt engineering and RAG approaches are insufficient.
- Engineer and version production prompts, integrating prompt and context management directly into the application layer.
- Integrate LLM APIs from providers such as OpenAI, Anthropic, and Azure OpenAI, as well as open-source model endpoints, ensuring proper authentication, rate-limiting, and cost controls.
- Instrument applications for evaluation by implementing output logging, quality scoring, and human feedback loops.
- Optimize latency and token costs through caching, batching, and intelligent model routing strategies.
- Translate complex client business requirements into concrete GenAI feature specifications.
- Communicate technical trade-offs related to cost, latency, and accuracy clearly to non-technical product stakeholders.
- Collaborate closely with AI architects and data scientists on shared components and solutions.
- Document architecture and prompt design decisions thoroughly to ensure maintainability and smooth handoffs.
Required Qualifications
- 4 to 8 years of software engineering experience, with 1 to 3 years specifically focused on hands-on GenAI and LLM application development.
- Strong proficiency in Python and experience with orchestration frameworks such as LangChain or LlamaIndex.
- Familiarity with vector databases and embedding strategies, including Pinecone, Weaviate, and pgvector, along with knowledge-graph or graph-database tools like Neo4j where applicable.
- Deep understanding of LLM failure modes such as hallucination, context-window limitations, and cost blowup, with the ability to design effective mitigations.
- Experience with model fine-tuning techniques like LoRA and QLoRA, as well as evaluation harnesses.
- Hands-on experience with enterprise GenAI and agentic platforms including Microsoft Azure AI Foundry, AWS Bedrock (including Strands Agents SDK), and Google Vertex AI.
- Knowledge of open-source frameworks such as LangChain and LlamaIndex is advantageous, especially in environments without mandated platforms.
- Proven skills in API design and integration, including managing authentication, rate limiting, and streaming responses.
- Familiarity with prompt-versioning and LLMOps tools such as LangSmith or Weights & Biases.
- Strong technical writing skills, capable of documenting complex RAG architectures for non-technical stakeholders.
- Comfortable working directly with client engineers during embedded delivery engagements.
- Collaborative mindset, able to work effectively with architects, data scientists, and QA teams without needing fully specified requirements.
This role offers the opportunity to work at the forefront of generative AI technology, contributing to impactful applications that serve enterprise clients. The successful candidate will be part of a dynamic team focused on innovation, quality, and delivering scalable AI solutions.