AI Security Engineer: What the Job Involves and How to Get Hired
What AI security means, the attacks and controls behind it, the skills, tools and projects employers look for, and how to get your CV past AI screening.
Demand for AI security engineers is growing because most companies now run chatbots, copilots or autonomous agents that no one has threat-modelled or tested. Job ads describe the work in very different ways, and many ask for several years of security experience. This guide explains what AI security covers, the attacks you should expect, how they are stopped, what recruiters' software looks for in your CV, and how to build proof of skill if you are new to the field. The industry is changing quickly and some of the concepts discussed here may be outdated within a few years, but AI security is likely to stay relevant for as long as we use this technology.
Two meanings of "AI security"
AI for security uses AI to do security work faster: triaging alerts, drafting detection rules, scanning code, running AI-powered vulnerability scans and even AI-powered penetration tests. These are now the norm, and many internal security teams are adopting or building such capabilities. SOC, detection and application security teams usually own and operate them.
Security for AI protects the AI system itself: the model, its training and retrieval data, its prompts, the tools it can call and the infrastructure that serves it. Application security, cloud security or a dedicated AI security team owns it. I do not expect most companies to build dedicated AI security teams; the work will more likely be absorbed into traditional security teams.
Most security practitioners do both, for example by using AI tools to test AI applications. Job ads often mix the two, so read the list of responsibilities and the title carefully.
Three kinds of roles
Across ads from large tech companies, banks and consultancies, the AI security work falls into three groups.
- AI in security operations. Employers want experienced practitioners in detection engineering, threat intelligence, incident response or cloud security, often with eight or more years behind them. They need people whose expertise improves what AI tools produce, so good domain security knowledge matters more than machine learning knowledge.
- Securing AI systems. You threat model LLM applications, test for prompt injection, protect model access and build guardrails. Typical titles are AI Security Engineer, Security Engineer (AI) and Application Security Engineer with an AI scope. Banks, consultancies and platform teams hire heavily here.
- Research. You build security capabilities into models, for instance reinforcement learning environments for security tasks. These roles combine machine learning with security expertise, are rare, and often favour a PhD or a strong research record.
Other titles you will meet include AI red teamer, AI governance, risk and compliance specialist, AI platform security engineer and agentic AI security specialist. Some ads only say "Security Engineer" and mention LLMs or agents in the description.
What AI security means in practice
An AI system exposes more attack surface than a conventional application. Attackers can influence the user prompt, any document, web page or email the model reads, the retrieval pipeline and vector database, the tools and plugins the model can call, its memory, other agents, and whatever consumes its output. Two properties make this hard. A model cannot reliably separate instructions from data, and it can answer the same input differently each time.
The main attacks
- Prompt injection. Direct injection overrides the system's instructions through user input. Indirect injection hides instructions in content the model processes, such as a web page or an email, and is the greater danger for agents because the user never sees it.
- Jailbreaking. Crafted input bypasses the model's safety rules.
- Sensitive data disclosure and system prompt leakage. The model reveals customer data, credentials or its own hidden instructions.
- Data and model poisoning. Tampered training, fine-tuning or retrieval data makes the model give wrong answers or carry a hidden backdoor.
- Supply chain compromise. Malicious model files, compromised libraries, plugins or unvetted MCP servers enter the system.
- Excessive agency. An agent holds more permissions than its task needs, so a hijacked prompt turns into a sent email, a deleted file or a moved payment.
- Improper output handling. Model output reaches a browser, shell or database unchecked and causes XSS, SQL injection or command execution.
- Vector and embedding weaknesses. Retrieval systems leak documents across users or tenants, or serve poisoned content.
- Unbounded consumption and model theft. Expensive requests inflate costs or deny service, and high-volume querying can copy a model's behaviour.
The OWASP Top 10 for LLM Applications and MITRE ATLAS catalogue these attacks and are the two frameworks interviewers expect you to know.
How these attacks are prevented
Controls fall into two families. Probabilistic controls are themselves models or instructions: system prompt rules, guardrail classifiers, LLM-based filters. They lower the success rate of attacks, but a determined attacker can still get through sometimes. Deterministic controls are enforced in ordinary code outside the model and give the same result for the same input every time. They form the foundation, because they hold even when the model has been fully manipulated.
Common deterministic controls:
- Scoped, short-lived credentials for every agent and user, following least privilege.
- Allowlists for the tools, domains and actions an agent may use, with all other outbound traffic blocked.
- Strict schema validation of model output before anything uses it, plus parameterised queries and output encoding.
- Human approval for irreversible actions such as payments, deletions and external emails.
- Sandboxed code execution without network access. Agents can still find ways out of a sandbox, so combine it with network restrictions and monitoring.
- Access control applied at retrieval time, so the model only sees documents the requesting user may read.
- Rate limits, token budgets and input size limits.
- Pinned and verified model versions, signed artefacts, dependency scanning and vetting of every MCP server.
- Logging of every prompt, tool call and response for detection and investigation.
The working principle is to assume the model can be manipulated and to make sure a manipulated model cannot do serious damage. Probabilistic guardrails are added on top to catch most attempts early, but the deterministic controls decide what a successful attack can actually do.
Security testing needs a different approach too. Because outputs vary, a single run proves little. Run each attack many times and report a success rate.
AI gateways: a control layer in front of the model
An AI gateway extends the authentication, routing and rate limiting of an ordinary API gateway to AI traffic. It sits in front of one or more model providers and adds model-specific controls: prompt and response inspection, redaction of personal data, token-based limits and cost budgets, model routing and fallback, audit logs and guardrails. Azure AI Foundry, AWS Bedrock and Google Vertex AI are platforms that host and serve models, and a gateway in front of them applies the same controls across every model and application.
Soft skills that decide interviews
Soft skills are the communication and judgement skills that let you turn a technical finding into a decision someone else can act on. AI security teams include researchers, product managers and lawyers who do not share your vocabulary, so interviewers test how clearly you explain risk and how practical your recommendations are. The following skills come up most often.
- Explain risk in business terms. "The model is vulnerable to prompt injection" means little to a product owner. "An attacker could make the chatbot reveal customer data, and this is the fix" gets a decision.
- Bring the fix with the finding. State the problem, the remedy and the effort required.
- Rate risk by likelihood and impact. Low, medium, high and critical let the business decide what to fix first.
- Understand how the company earns money. If you know who the customers are and which failures would cost the most, you can tell which AI systems and which data need the strongest protection. Research its business-critical systems, for example in its annual report or product pages.
- Work across teams. AI security engineers spend much of their time with researchers, product managers and legal teams.
- Conflict resolution and teamwork. AI security is a fast-moving field with tight deadlines and disagreements over priorities, so stressful situations are likely. Read articles or watch videos on handling them professionally.
Technical skills: attack, defence and the tools to practise them
The technical base for AI security is ordinary security knowledge plus an understanding of how AI applications are built: where prompts, data, tools and permissions meet. Each skill below can be practised at home or in a lab.
Languages
Python is the default for AI tooling, test harnesses and scripting. Go and Rust are becoming popular for security tooling and network services because they compile to fast native binaries. Rust's memory safety removes a whole class of vulnerabilities, and Go's simplicity and concurrency suit proxies and gateways. Basic Bash and SQL cover automation and data access.
Offensive skills
- Prompt injection and jailbreak testing. Start by hand. Lakera's Gandalf game and the large language model labs in PortSwigger's Web Security Academy teach direct and indirect injection, system prompt extraction and filter bypass step by step. Once you understand the techniques, automate them: garak from NVIDIA fires hundreds of probe prompts at a model, PyRIT from Microsoft orchestrates multi-turn attack scenarios, and promptfoo runs red-team test suites that you can add to a CI pipeline.
- Web and API testing. AI applications are web applications with APIs, so the classic flaws still apply: a chatbot that builds SQL from model output is open to SQL injection, and an insecure plugin endpoint can leak data through broken access control. Learn Burp Suite to intercept and modify the requests between the browser, the application and the model API, and OWASP ZAP for automated scanning. PortSwigger's free Web Security Academy covers SQL injection, SSRF, access control and API testing, and completing its labs is some of the best preparation available.
- Agent and MCP attacks. Run an intentionally vulnerable agent in Docker, such as Damn Vulnerable LLM Agent or Damn Vulnerable MCP Server, and try three concrete attacks: hide instructions in a web page or email the agent summarises, trick it into calling a tool with attacker-chosen arguments such as sending data to an external address, and write false facts into its memory so that later sessions act on them. Record which attacks succeeded and what stopped the rest.
- Cloud and container attacks. Use a deliberately vulnerable cloud lab such as CloudGoat for AWS to exploit an over-permissive IAM role, a public storage bucket and an exposed metadata service, the same weaknesses that expose model endpoints and training data. For containers, practise escaping a Docker container that has a mounted Docker socket or privileged mode, because agents often execute code inside containers.
Defensive skills
- Threat modelling. Threat modelling is learned by doing it on a real system. Take a simple chatbot that answers questions from your own documents and draw a data flow diagram with draw.io or OWASP Threat Dragon. Mark the trust boundaries: user to application, application to model API, application to vector database, and agent to tools. Apply the STRIDE categories (spoofing, tampering, repudiation, information disclosure, denial of service, elevation of privilege) to each element, then map every threat to the OWASP Top 10 for LLM Applications and MITRE ATLAS. Rate each threat by likelihood and impact and write down one control for each high-rated item. The OWASP Threat Modeling Cheat Sheet and Adam Shostack's book Threat Modeling: Designing for Security explain the method, and repeating the exercise on every project you build is the fastest way to improve.
- Deterministic guardrails. Build a small agent with an email-sending tool, then enforce limits in code instead of in the prompt. Validate every tool call against a JSON Schema or Pydantic model, allow only approved recipient domains, give the agent an API key limited to the one action it needs, and require human approval for anything irreversible. Open Policy Agent lets you write these rules as policies outside the application. Finally run a prompt injection attack against the agent and confirm that the code blocks it even when the model complies.
- Detection and logging. Record every prompt, response and tool call as structured JSON, using a tracing tool such as Langfuse or OpenTelemetry, and send the logs to the Elastic stack or Splunk. Then write detection rules, in the portable Sigma format or in your platform's own query language, for patterns such as a sudden rise in refusals, a tool call to a domain never seen before, unusually long prompts, and responses that contain fragments of the system prompt. Test each rule by replaying your own attacks from the offensive labs.
- Code and cloud scanning. Add a GitHub Actions pipeline to your own repository that runs Semgrep or CodeQL on the code, Trivy on the container image and Checkov on the Terraform files, and set it to fail the build on high-severity findings. Prowler audits an AWS account for misconfigurations. Fixing the findings in your own project gives you something concrete to describe in an interview.
- Web application firewall (WAF) configuration. A WAF filters HTTP traffic in front of an application. It blocks classic attacks such as SQL injection and cross-site scripting and can rate limit abusive clients to protect a costly model endpoint, but it does not understand what a prompt means, so it is one layer among several. Practise by running ModSecurity with the OWASP Core Rule Set in Docker in front of a vulnerable application such as DVWA, tune the rules to remove false positives and confirm the results with Burp Suite. cyberlearner.org has practical defence labs that include configuring a WAF. Cloud equivalents such as AWS WAF and Cloudflare use the same concepts.
Where to practise
Hack The Box and TryHackMe offer guided labs from beginner to advanced. cyberlearner.org is my own free CTF portal for hands-on challenges. At home, you can run open models locally with Ollama and Docker to build and attack your own test applications.
Your CV will probably meet an AI first
There is a high chance that the first stage of recruitment is automated. Applicant tracking systems and AI screening tools scan your CV for years of experience, keywords, and relevant projects or research before a person sees it. The tips below help, but stay honest in your CV:
- Use the terms from the job ad wherever they are true: prompt injection, LLM security, agent security, threat modelling, OWASP LLM Top 10, your cloud platform.
- State your years of experience and employment dates clearly.
- Put AI security projects and research near the top, with links.
- Use a plain single-column layout without tables or graphics, which parsers often misread.
- Give measurable results instead of task lists.
A human reads the CV after the software does, so avoid stuffing keywords you cannot discuss.
Projects that prove you can do the work
If you lack the years of experience, projects supply the evidence. Publish them on GitHub with honest write-ups, and list them on your CV.
- Build a vulnerable LLM application, exploit it, then fix it. Indirect prompt injection through a retrieval pipeline is a good first target, and the write-up should cover how you tested the fix.
- Attack intentionally vulnerable agents and MCP servers. Open-source projects exist for this, and a write-up of what you broke is useful evidence.
- Build a prompt injection detector. Even an imperfect one shows you understand the problem.
- Reproduce a published attack or defence. Pick a recent paper or vendor advisory, rebuild the attack or defence against a test application of your own, and record the exact setup, the number of runs and the success rate. Report which claims held up, which did not and what you changed to make it work. A write-up that shows where a published result fails is more convincing to employers than one that confirms it.
- Design a method for measuring how consistent a model's safeguards are. Choose one behaviour a model should refuse, such as revealing a hidden system prompt, and write several differently worded prompts that ask for it. Send each prompt dozens of times with the same settings, count how often the model refuses, and report the refusal rate per prompt. A safeguard that blocks a prompt in 95 of 100 runs still fails in 5, and the prompts with the lowest rates show where it is weakest. Because the same input can produce different answers, a single test proves little, and repeated measurement gives a defensible number. Publish the prompts, settings and results so others can reproduce them.
- Compare an LLM-based code scanner with a traditional one. Traditional scanners to use include Semgrep, CodeQL, SonarQube and Bandit for Python. Run both approaches on the same deliberately vulnerable codebase, then document where each finds issues the other misses, where each produces false positives, and where the LLM invents vulnerabilities that do not exist.
- Test sandbox escapes. This is an emerging research area. Recent publications show that agents can find ways out of sandboxes, containers and even virtual machines, and companies running agents are likely to invest in preventing it. You can reproduce these scenarios at home with container images or a cloud lab, then add your own protection, such as stricter permissions, network restrictions or monitoring, and test whether it holds.
Certifications
Certifications are largely a tick-box exercise, but an important one: many employers and screening tools filter on them before anyone reads your projects, so skipping them can cost you interviews. They do not prove you can do the work, so treat them as an entry requirement and not as your main evidence. A sensible set combines one general or hands-on security certification, one cloud security certification and one AI security certification. For the first, CISSP suits senior roles, OSCP suits offensive work and BSCP (Burp Suite Certified Practitioner) suits web application testing. For the second, choose the platform your target employers use, for example Microsoft Azure security engineer certification or AWS Certified Security – Specialty. For the third, options include OffSec's AI red teaming certification, GIAC's AI-focused certifications and the entry-level ISC2 and CompTIA AI security certificates.
What the roles pay
Advertised UK salaries in 2026 mostly fall between £75,000 and £110,000 for AI security engineers, with lead roles at banks reaching £130,000 and top AI labs paying considerably more. Entry-level roles start around £60,000, and London contract roles pay roughly £500 to £600 per day. Advertised US ranges run from about $125,000 to $235,000. Check current ads for your market, since ranges vary widely by employer and seniority.
Where to apply and how to get noticed
Banks, hospitals, retailers and consultancies all deploy AI and need people to secure it, so apply beyond AI companies. OWASP chapters, BSides and local security meetups put you in front of the small community that does the hiring. Following new adversarial research each month and writing about your own attempts to reproduce it keeps your knowledge current and visible.
Practise for interviews
The interview preparation platform cyberole.org has AI security interview questions built from real job descriptions, labs for securing LLM applications and testing agents, and incident scenarios for practising a response to an AI breach. You can start preparing for free today.
Conclusion
Applying for jobs is stressful right now, and AI makes it harder: recruiters receive far more applications than they can read, so software often decides who gets through. This guide covered the two meanings of AI security, the main attacks and the deterministic controls that stop them, the technical and soft skills employers test, how to get your CV past automated screening, and the projects and certifications that prove your ability.
Projects, experience and research are what set you apart from the other applicants, so build them, publish them and apply widely. Keep going after rejections. Organisations keep adopting AI and need people to secure it, which makes this a strong time for security professionals and aspiring students to start on the career ladder.
References
Frameworks and methods
Offensive tools and labs
- Lakera Gandalf
- PortSwigger Web Security Academy and its LLM attacks labs
- garak, PyRIT and promptfoo
- Damn Vulnerable LLM Agent and Damn Vulnerable MCP Server
- CloudGoat
Defensive tools
- Open Policy Agent, Langfuse, OpenTelemetry and Sigma
- Semgrep, CodeQL, Trivy, Checkov and Prowler
- ModSecurity, OWASP Core Rule Set and DVWA
Practice platforms
Certifications
- CISSP (ISC2)
- OSCP (OffSec)
- Burp Suite Certified Practitioner (PortSwigger)
- AWS Certified Security – Specialty
- GIAC certifications
- OffSec, ISC2 and CompTIA for the AI security certificates
Books
- Adam Shostack, Threat Modeling: Designing for Security (Wiley, 2014)