Role: Agentic AI Engineer - Cloud Infrastructure Automation
Location: US, Remote
Contract
Role Summary
We are seeking an Agentic AI Engineer to design and deliver AI agents that operate against live cloud infrastructure. This is a hands-on build role: you will stand up agent workflows that interpret infrastructure state and telemetry, generate and validate Infrastructure-as-Code, and execute controlled remediation and provisioning actions under human-in-the-loop approval.
This is not a research or prototyping role. The expectation is working, governed automation deployed into a real environment within the engagement window.
Key Responsibilities
· Build and deploy agentic workflows using an LLM orchestration framework (LangGraph, CrewAI, AutoGen, Semantic Kernel, or cloud-native equivalents such as Bedrock Agents / Azure AI Agent Service)
· Design and implement tool/function-calling layers that expose infrastructure operations to agents safely — including MCP-based tooling where applicable
· Generate, review and validate Terraform modules programmatically; integrate policy-as-code checks (OPA/Sentinel, tflint, checkov) into agent output paths
· Integrate agents with observability and ITSM systems (DataDog, ServiceNow, PagerDuty, Jira) for alert triage and automated ticket-to-action flows
· Implement guardrails: least-privilege execution roles, approval gates, dry-run/plan-only modes, audit logging of every agent action
· Build evaluation and regression harnesses for agent behaviour (accuracy, tool-selection correctness, failure containment)
· Wire agent workflows into existing CI/CD pipelines
· Document architecture and hand over to the client's platform team at engagement close
Required Skills
Agentic AI / LLM
· 2+ years building production LLM applications, including at least one deployed multi-step agentic system
· Strong command of one orchestration framework (LangGraph/LangChain, CrewAI, AutoGen, Semantic Kernel) plus tool/function calling and structured output
· Practical experience with agent guardrails, retries, failure handling and human-in-the-loop approval design
· RAG fundamentals: embeddings, vector stores (pgvector, OpenSearch, Pinecone), retrieval quality tuning
· Agent evaluation and observability (LangSmith, Langfuse, Ragas, or equivalent custom harnesses)
Infrastructure / Platform
· 5+ years in cloud infrastructure, platform or DevOps engineering
· Deep Terraform: module design, state and remote backends, workspaces, drift handling; Terragrunt or Terraform Cloud/Enterprise a plus
· Policy-as-code: OPA/Rego, Sentinel, or equivalent
· Cloud depth in at least one of AWS / Azure / OCI, including IAM and least-privilege design
· CI/CD: GitHub Actions, Azure DevOps, GitLab CI or Jenkins
· Containers and orchestration (Docker, Kubernetes)
· Observability tooling - DataDog strongly preferred
Engineering
· Expert-level Python (async, API integration, testing)
· Secrets management (Vault, AWS Secrets Manager, Azure Key Vault)
· Git-based workflows and code review discipline
Nice to Have
· Prior experience building agents that take write actions against production infrastructure
· Model Context Protocol (MCP) server/client implementation
· FinOps tooling exposure (Cloudability, Apptio)
· Oracle Cloud Infrastructure (OCI)
· Healthcare or regulated-industry environment experience