NeuralSurge AI is seeking a Senior AI & Machine Learning Engineer to lead the technical development and ownership of our end-to-end intelligence pipelines. This role involves working across our proprietary product platforms as well as client services, bridging experimental data science models with customer-facing full-stack applications. You will be responsible for deploying and optimizing open-weight foundation models on dedicated GPU infrastructure, enhancing inference performance, and collaborating closely with our MERN stack engineering team to deliver scalable, high-performance AI solutions.
Key Responsibilities
- End-to-End Pipeline Execution: Architect, deploy, monitor, and scale production AI/ML pipelines, including data ingestion, custom chunking, embedding generation, vector indexing, inference orchestration, and evaluation loops.
- Inference Serving & Runtimes: Deploy and optimize open-source foundation models such as Llama, Qwen, Mistral, and Gemma using high-throughput runtimes like vLLM, Triton Inference Server, Ollama, or TensorRT-LLM on GPU infrastructure.
- Pipeline Engineering: Design and build custom retrieval systems, rerankers, semantic query routers, and agentic reasoning loops using raw Python and vector databases, avoiding brittle black-box abstractions.
- Model Fine-Tuning & Weight Adaptation: Lead parameter-efficient fine-tuning workflows including PEFT, LoRA, QLoRA, Unsloth, and TRL, adapting open-source models to meet domain-specific enterprise needs.
- MERN Stack & Microservices Integration: Work closely with the MERN engineering team to expose low-latency, asynchronous FastAPI endpoints, streaming protocols such as WebSockets and Server-Sent Events, and background worker contracts using Redis, Celery, or BullMQ.
- Dual-Track Delivery:
- *Product Track:* Build and strengthen the core intelligence layers powering NeuralSurge’s proprietary SaaS platforms.
- *Services Track:* Architect secure, compliant, and sovereign/air-gapped AI pipelines tailored for enterprise client engagements.
- Performance & Cost Optimization: Profile inference latency, token economics, batching, and GPU VRAM usage using quantization techniques like AWQ, GPTQ, and GGUF to maximize throughput while minimizing compute overhead.
- LLMOps & Quality Engineering: Implement automated evaluation suites such as RAGAS, DeepEval, and LLM-as-judge, along with latency tracing and deterministic guardrails to prevent hallucination, data leakage, and model drift.
Required Qualifications
- Bachelor’s degree in Computer Science or a related field with over 3 years of professional engineering experience delivering production-grade Python backend services or applied AI/ML systems.
- Proven experience deploying and maintaining AI/ML pipelines in production environments.
- Strong proficiency in Python and familiarity with vector databases and AI model fine-tuning techniques.
- Experience working in collaborative, cross-functional teams, particularly with full-stack engineers.
Preferred Qualifications
- Practical exposure to multimodal architectures including Computer Vision (e.g., YOLO), voice agents (e.g., Whisper), or diffusion models.
- Experience designing and deploying air-gapped or on-premise sovereign AI infrastructure for enterprise clients.
- Contributions to open-source machine learning tooling, custom frameworks, or published research in relevant AI/ML domains.
NeuralSurge AI offers the opportunity to work onsite in Lahore, Pakistan (Etihad Town, Phase-1) in a full-time capacity, contributing to cutting-edge AI infrastructure and enterprise-grade solutions. This role is ideal for engineers passionate about pushing the boundaries of AI deployment and performance optimization in real-world applications.