How do you create an AI roleplay? A step-by-step guide to scenario design, personas, scoring rubrics and testing (plus how to pick the right platform).

Most corporate training is something people sit through. In a 2026 Training Industry survey of more than 1,200 learners, 67% admitted to multitasking their way through it. AI roleplay breaks that pattern for a simple reason: you can't half-listen to someone who is looking right at you, waiting for you to answer.
That's what makes it worth building. Instead of watching a video about a tough conversation, your people get to have one — as many times as they need. Teams practicing this way see 95% training effectiveness, twice the knowledge retention of text-based learning, and 3X the engagement. 90% say it's less stressful than roleplaying with a colleague, which matters more than it sounds.
A roleplay is a rehearsal, and rehearsals work when three things are true. The thing you're rehearsing is specific, the other person behaves like a real person, and somebody tells you what to fix afterward. Get those right and you'll build something your team comes back to.
This guide walks through the whole build, start to finish. It's platform-agnostic — you can follow it inside UneeQ's Immersive Training Platform, inside a competitor, or with a plain LLM and a lot of patience. We build this software for a living, so we'll be upfront about which parts a platform handles for you and which parts are still your job.
An AI roleplay is a practice conversation between a learner and an AI character that responds in real time. The learner talks; the AI listens, reacts, pushes back, and stays in character; and afterward the learner gets scored and coached on how they handled it.
It's often used for the conversations that are expensive to get wrong: discovery calls, escalations, performance reviews, negotiations, job interviews. Unlike a video course or a quiz, nothing is scripted. The conversation goes wherever the learner takes it.
For the full picture, read our guide: what is AI roleplay.
You need four things as the foundation for a good AI roleplay initiative.
If you can't fill in all four, take a pause. You'll need clarity to be able to move forward effectively.
OK, let's take it step by step. But don't worry, some of these steps take a matter of minutes. Alternatively, you can watch the below video on how you can create a scenario quickly on UneeQ's Immersive Training Platform.
A good AI roleplay starts with a simple question: What should someone be better at after practicing this conversation?
Everything else (the scenario, scoring rubric, difficulty, persona behavior, and feedback) should work backwards from that.
Specificity is the name of the game here. A vague brief creates a vague conversation, and a vague conversation teaches nobody much.
Here's the difference in practice.
Weak brief: "Talk down a frustrated Head of IT before they switch vendors."
Strong brief: "The participant is an Enterprise Account Manager. They are talking to Emily White, the Head of IT, who is frustrated and withdrawn, feeling unheard. A critical system integration project is significantly behind schedule, causing major operational issues for the client, and Emily has been unresponsive to recent communications. The client is actively considering switching vendors, putting the entire account at risk. The participant's goal is to understand the full depth of the issue, rebuild trust, and propose a clear, actionable path to resolution. A successful outcome is Emily fully expressing her concerns and agreeing to work collaboratively on a recovery plan, feeling confident in the EAM's commitment. Emily's challenging behavior is going quiet, offering minimal, one-word responses, forcing the EAM to ask very direct, probing questions to get any information."
The second one has a starting emotional temperature, a history, a specific ask, and a hard constraint on the learner. That's enough for an AI roleplay platform to behave consistently and enough for you to score the result.
Every scenario brief should answer five questions:

Sales — sales training
"The participant is an account executive making a first discovery call with Jordan Mehta, a VP of Operations at a 500-person logistics company who agreed to the meeting but is skeptical that yet another software tool will help. The goal is to uncover Jordan's biggest operational pain points, qualify budget and timeline, and earn a follow-up demo. Jordan is friendly but guarded, pushes back on vague claims, and will end the call early if the rep launches into a generic pitch. A successful outcome is the rep asking sharp, open questions, surfacing at least two concrete pain points, and securing a scheduled demo with the right stakeholders."
Customer service — customer service training
"The participant is a customer support agent. They are talking to Marcus Bell, a small business owner who is angry and convinced he has been overcharged. Marcus stopped using the software eight months ago but never cancelled his subscription, and has now noticed eight monthly charges on his statement. He believes the company should have flagged the inactivity and refunded all of it. The account is worth little, but a public complaint would cost more than the refund. The participant's goal is to acknowledge the frustration without accepting blame the company doesn't own, explain the charges clearly, and land on a resolution Marcus accepts. A successful outcome is Marcus agreeing to an immediate cancellation plus a goodwill refund of the two most recent months, understanding why the remaining charges stand. Marcus's challenging behavior is talking over the agent, returning repeatedly to "I never even logged in," and threatening a chargeback and a public review."
Leadership — leadership training
"The participant is a manager meeting with Chloe, a Data Analyst who is disengaged and apathetic. Chloe consistently meets basic requirements but shows no initiative, avoids new challenges, and her lack of engagement is starting to affect team dynamics. The stakes are Chloe's long-term potential, team innovation, and overall department productivity. The participant's goal is to re-engage Chloe, identify underlying reasons for her disengagement, and collaboratively explore development opportunities that align with her interests and company needs. A successful outcome involves Chloe expressing renewed interest, agreeing to explore specific development paths, and committing to taking more initiative. Chloe's challenging behavior is going quiet, offering minimal, one-word responses, making it hard to gauge her perspective or commitment."

Let's talk about your AI roleplay scoring rubric. The rubric is where your methodology lives, which will drive what comes up in your 3D Analytics and AI coaching within Immersive Training Platform.
These can be ultimate aims of the call, such as to "agree to a follow-up call next week", but can also include the types of behaviors you want to coach. It can also include things said on the call that will cause an automatic fail.
For instance: "A successful outcome is the rep asking sharp, open questions, surfacing at least two concrete pain points, and securing a scheduled demo with the right stakeholders". That's enough information for 3D Analytics to track during roleplay, and for your AI coach to provide feedback on.
There's also an advantage to focusing on this specific type of feedback. It helps managers discover the bahavios, characteristics, and habits of their staff members, and coach them in a particular way that doesn't focus judgments on who they are but on what they do.
There's evidence behind that distinction. Kluger and DeNisi's meta-analysis of 607 effect sizes across 23,663 observations found that a manager's feedback improves trainee performance overall, but more than a third of the interventions actually reduced it, especially when the attention of the feedback shifted away from the task and toward the self.
A manager telling a direct report "you're not a very good listener" is an unhelpful judgment about the person, even when it's true. Whereas "you interrupted the customer twice while she was explaining the billing history" describes a behavior the learner can change.
Feedback works better when it focuses on what someone did, rather than what kind of person they are, which is something to remember when creating the scoring rubric in your AI roleplay.
Difficulty is a big part of strong scenario design too. Too easy and the simulation manufactures false confidence; too hard and you might not be replicating an actual winnable scenario your staff will face in real life.
You can increase difficulty by changing how emotionally charged the persona is, how much information the learner receives beforehand, how cooperative the other person is, how much time they have, or how firmly objections are defended.
A good trick is to increase difficulty across a learning journey rather than throwing everything into one nightmare scenario.
A learner might first practice the structure of the conversation with a relatively cooperative persona. Next, they encounter realistic resistance. Finally, they face someone skeptical, time-poor, frustrated, or unwilling to volunteer information. That replicates progression rather than punishment, and it simulates the many types of people your staff will encounter in real life.
Keep an eye on your team-wide training analytics. If 99% of learners are failing the endorsement criteria in a scenario, it's probably a sign you need to lower the difficulty, make the objectives clearer, or make other tweaks that mean more learners have a fair chance of success.
A thin persona produces an agreeable AI no one wants to roleplay with because it doesn't feel real, and therefore fails the realism test. So let's avoid some of the common slip-ups that lead to messy personas that don't fit the bill.
On Immersive Training Platform, you'll notice that persona creation is a specific step in the process. Here you choose the look of the digital human you'll roleplay with.
Next, you choose their personality type, which determines how they'll respond in each session. There are a number of pre-built personality types to choose from (friendly, indifferent, skeptical, confronting), or you can create a custom personality to match the persona you have in mind.
The platform will present a template on what to include to create a realistic personality type, which includes:

A hill we'll die on is that no one needs to roleplay the conversations that go well – the ones that don't test a learner's ability to stay composed, on-brief, and professional.
So don't fall into the trap of making the AI too agreeable. Here are some tips to avoid that kind of outcome:
Then test it by being deliberately terrible (see Step 3). If you can waffle your way to a win, the persona isn't finished.

Before you rollout your new roleplay scenario, you have the option to test it via the 'Try it now' button.
A good way to test is to play through it once attempting to get a good score, and again playing it badly on purpose. If the terrible run still passes, your scoring rubric will need tweaking. You can also launch new roleplays to smaller groups, set up in your admin portal, who can battle-test it before the whole company starts to use it irl.
From this small test, you can look at metrics like:
As well as looking at the data, you should also ask the small cohort for direct feedback. Did the persona feel realistic and similar to what they face in the field? If so, great; if not, you can easily refine your persona.
Keep an eye on that team-level data – you might even want to set a reminder to revisit training insights every month or quarter. Doing so might just unearth some hidden strengths and weaknesses in your team.
When 3D Analytics shows that 70% of your sales reps lose control of the conversation at the same objection, congrats, you've found next month's team meeting. The sales leaders will love you for it!
After launch, it's easy to edit a roleplay scenario should you need to. You can edit directly or choose 'refine with AI' within Immersive Training Platform. The latter will allow you to describe the changes you want, rather than manually editing line by line.
The beauty of an AI roleplay platform like ours is the ease and speed at which you can iterate your existing scenarios, so if your products, sales motion, value proposition, or company policy changes, your sessions can change too with just a couple of minutes of tweaking.
Roughly 20 to 40 minutes for the first scenario, and 5 to 10 minutes for each one after that, assuming you already know what conversation you want practiced.
Building a scripted branching roleplay used to be a months-long project involving storyboards, content creation, and a budget approval. Now it's a paragraph of plain English and a few clarifying questions.
Start from real conversations your people are already having, then build the smallest set that covers your highest-risk moments.
The practical sequence for an organization, rather than for one scenario:
Yes, and for a single person practicing a single conversation it works well. A well-written persona prompt in ChatGPT or Claude will give you a useful sparring partner in about five minutes, for free(ish).
Where it stops working is consistency and scale. Specifically:
If you're one person prepping for one difficult conversation next Tuesday, open ChatGPT and go for it. If you're responsible for 200 people and someone is going to ask you next quarter whether the training worked, you need consistent scoring and a record – you need a purpose-built roleplay platform.
Yes, and it's the fastest path to a scenario that feels real, because you're not relying on imagination, but true-to-life events that happen in your teams. We focused above on creating a scenario from calls that have cost you money, but you can also use examples that have gone well, too.
How to do it:
That's it. Immersive Training Platform will take your call recording and use it to build a roleplay scenario that your whole team can practice with.
To assess the type of roleplay platform you nee, pick based on where the conversation actually happens in real life. That's a better way to reach a productive decision before you start comparing feature lists.
If your team works the phones, a voice-only tool covers it. Hyperbound, Second Nature, and Yoodli all work this way. They're usually cheaper, and there's no sense paying for video fidelity nobody's going to use.
If the conversation happens on a video call or in-person, you need a platform that simulates a face-to-face interaction. Did they hold eye contact while delivering bad news? Did they clock the customer folding their arms when the price came up? Did they keep going when the character got difficult, or fold at the first push-back? None of that shows up in an audio file, and you can't coach what you can't see. UneeQ's Immersive Training Platform is built for in-person roleplay, as are alternatives like Virti and Mursion.
Match the format to where the conversation happens, then check whether anyone needs to report on the results. These four questions should help you decide:
Six things separate the scenarios people learn from and the ones they click through:
That last pair is the argument for roleplay in one line. Confidence and competence come apart, and only one of them shows up on the call.
Now, the answer here is obviously multifaceted. But, let's be straight: you'll want to measure ROI, engagement rates, and effectiveness, among other team-specific metrics.
For reference, these are some of the tangible results we've achieved with enterprise L&D teams at UneeQ.
