About AI Platform
We are developing a shared AI platform for the business units of a multi-industry group in Vietnam. The platform runs on Kubernetes with LangGraph, FastAPI, and AWS Bedrock, using LangFuse for observability, and enables business teams to build and operate their own AI agents: attaching tools, loading internal documents, and orchestrating multiple agents that work together.
About the Role
We’re looking for a Senior Backend Engineer to design and build scalable, production-ready AI solutions and shared platform capabilities.
You’ll combine hands-on development with technical direction, making architectural decisions and establishing engineering standards. Your work will balance solution quality, performance, reliability, security, and cost, from early proof of concept through production and scale.
What You’ll Do
- Design end-to-end AI/GenAI architectures, selecting suitable foundation models, embedding models, vector databases, search engines, agent frameworks, and orchestration strategies.
- Build RAG, agentic AI, tool-calling, and AI workflows, defining the appropriate boundaries between deterministic execution and LLM-driven reasoning.
- Develop reusable platform capabilities, including Model Gateway, Prompt Management, RAG Pipelines, Knowledge Management, Agent Runtime, Tool Registry, Guardrails, Evaluation Framework, and AI Observability.
- Design and develop a platform that supports multiple teams, products, and use cases.
- Build production monitoring and observability, implement failure handling and graceful degradation, and optimize latency, token consumption, and inference cost.
- Benchmark models for specific use cases and improve prompts, context, retrieval, model selection, routing, caching, and inference.
- Design and implement evaluation frameworks, datasets, metrics, and production quality gates covering retrieval quality, answer relevance, factuality, hallucinations, task completion, and safety.
- Guide technical solutions, review architecture and code, establish engineering best practices, mentor AI Engineers, and proactively address technical risks.
- Partner with Product Owners and AI Specialists to assess feasibility and ROI, translate business requirements into architecture, and challenge use cases where AI may not be appropriate.
- Support solutions throughout their lifecycle: POC → MVP → Production → Scale.
What We’re Looking For
- A bachelor’s degree in Computer Science, Information Technology, or an equivalent qualification. Certifications are not required.
- 5+ years of experience in Software Engineering or AI Engineering, with proficiency in Python, Java, Go, or a comparable language.
- Practical experience building and deploying LLM/GenAI applications, with a strong understanding of LLM architecture and capabilities, prompt engineering, RAG, embeddings, vector search, agentic AI, tool calling, evaluation, and guardrails.
- Experience designing distributed systems, APIs, microservices, and cloud-native applications, supported by solid knowledge of databases, caching, and messaging.
- Experience designing production-grade AI systems and making architectural decisions based on business requirements and technical constraints.
- Experience with AWS, Azure, or GCP, along with Docker/Kubernetes, CI/CD, infrastructure as code, and production monitoring and observability.
- The ability to evaluate trade-offs between quality, latency, scalability, and cost, and determine when to use RAG, fine-tuning, agents, traditional ML, deterministic workflows, or LLM reasoning.
- Proficiency in Git branching and release strategies, and experience documenting architecture and decisions through tools such as Jira, Confluence, and ADRs.
- The ability to explain and defend architectural decisions to technical and non-technical stakeholders, make decisions with incomplete information, and clearly communicate assumptions.
- Strong English reading and writing skills, constructive communication, and ownership of systems throughout their production lifecycle.
Nice to have
Experience building enterprise AI platforms; using LangGraph, LangChain, or similar frameworks; developing agentic workflows; implementing model gateways, LLM routing, or multi-model architectures; building LLM evaluation and observability platforms; deploying open-source LLMs; working with AWS Bedrock, SageMaker, or other managed AI services; or operating in environments with demanding security, governance, and compliance requirements.
