Design and develop internal flagship data analytics systems, applications and APIs that allow engineers and analysts to retrieve, triage and analyse information more efficiently
Work with product managers, engineering managers and key stakeholders to deliver impactful solutions that meet our business needs
Manage enterprise system performance, reliability and sustainability through software quality control and optimisation of software products and technologies
...
AI Platform Engineering & Operations: Design, deploy and sustain enterprise-grade AI platforms and common AI services, including LLM hosting, inference services, embedding services and AI agent orchestration platforms.
Model Deployment & Optimisation: Deploy, optimise and manage AI models across on-premise and cloud environments. Improve model inference performance, throughput and resource utilisation through continuous monitoring, observability and optimisation techniques such as quantization, vLLM, batching etc.
GPU Infrastructure & Resource Management: Manage and optimise GPU infrastructure and compute resources to ensure efficient allocation, scalability and high utilisation across enterprise AI workloads. Support capacity planning, workload scheduling and performance monitoring for AI systems.
...
Design, develop, and implement highly scalable and reliable data pipelines and infrastructure.
Take full ownership of complex technical problems, from ideation and research to proposing and delivering end-to-end solutions, often involving ambiguous requirements.
Drive technical strategy and best practices for data ingestion, processing, storage, and consumption, ensuring data quality, compliance (e.g., data localization, privacy), and discoverability.
...
Agentic AI ArchitectureDesign and implement production-grade agentic systems using models such as Claude. Agent orchestration, tool calling, function calling, multi-agent architectures, planning and task decomposition, agent memory, context management, state machines and workflow engines, long-running agents, human-in-the-loop systems, autonomous execution, recovery and retry mechanisms, observability, and evaluation. You understand the difference between LLM → Agent → Workflow → Autonomous System, and when each is appropriate.
Agentic Loops & Self-OptimizationA major part of the role is building closed-loop systems: Goal → Plan → Execute → Observe → Evaluate → Learn → Re-plan → Execute. Systems that evaluate their own outputs, detect failed actions, identify root causes, adjust strategies, optimize prompts and tool selection, keep what works, roll back what does not, and improve over time. Reflection, critique, self-evaluation, feedback loops, reward signals, evaluation frameworks, automated experimentation, memory, retrieval, state management.
Claude / LLM EngineeringDeep practical experience with Claude/Anthropic APIs is highly desirable. Tool use, structured outputs, streaming, context management, prompt engineering, system prompts, long-context workflows, model routing, token optimization, latency and cost optimization, context compression, agent memory, LLM evaluation. Experience with other frontier models (OpenAI, Gemini, Llama or equivalent) is a plus.
...
Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesaving drugs, forecast supply chain disruptions, locate missing children, and more.
If you're a dreamer and builder who wants to craft the digital world of tomorrow alongside passionate visionaries, join us. We never stop striving for seamless, secure and intelligent collaboration solutions that drive productivity, insights, and innovation. Dream it, build it, and make an impact with AvePoint.