# Designing research for AI moderators: Lessons from running 100 interviews a week — Rebecca Klee — session 2026-08-26T00:05:00.000Z → 2026-08-26T00:35:00.000Z

_468 transcript lines · 79 slides_

## Transcript

The capabilities will change. The need for research judgment will last longer. So, Before the design, what does AI moderation actually offer? You've probably heard the pitch qualitative depth at survey scale. By a show of hands, who feels confident that AI can genuinely achieve depth at scale? Who is unsure? Who is skeptical? That makes sense because depth is doing quite a lot of work. Surveys give us reach but limited depth, Interviews give us depth, but limited reach. AI moderation appears to offer both a conversation every participant, repeated hundreds or thousands of times, But I don't think depth at scale is quite right. Seyda Bakshi, a quantitative researcher at OpenAI, calls this framing a fallacy because it treats methods as a single line with AI moderation finally giving us everything at once. And I AI moderator can ask what changed, why something stopped feeling worthwhile, someone someone did instead, That is meaningfully deeper than an open text response. But more probing produces more words. Not necessarily everything we associate with a depth interview. A moderator can't yet observe behavior, notice hesitation, or investigate the gap between what someone says and what they do. Nor can it reliably uncover reasons people don't understand or aren't willing or able to articulate. And more sessions don't provide a valid measure of prevalence. If thirty out of a hundred participants mentioned something, that of course doesn't mean that thirty percent of the population experience it. AI moderation offers a particular kind of depth with particular structural limits. A more useful way to think about this is that different methods answer different parts of the research problem. Surveys tell us how many people experience something and how that varies between groups, but are largely limited to reasons we already know to ask about. Human moderated interviews reconstruct what happened, how someone behaved, and how an experience felt. But their scale limits claims about prevalence, I Open text can surface reasons that we didn't anticipate, but tend to stop at a participant's first answer. So AI moderation. Fits between them. It can take a volunteered reason and ask what sits underneath it. And repeat that across many participants. Bakshi calls this probed language at machine scale. Now, In my program, that language was spoken rather than typed. But the capability was the same. So rather than asking whether AI gives us depth at scale, I now ask what kind of evidence is this method So That question is much more useful when deciding whether to use a it. So where might AI moderation actually earn its place? The first use case is recurring or time critical use research. My program moves from a focused question to completed sessions and an insight report each week. Because interviews run-in parallel, evidence can inform a decision without waiting weeks for sessions to be conducted sequin That speed is valuable when it supports a real decision. Not simply because faster is always better. The second use case is broader qualitative comparison, exploring how and why experiences differ across customer groups or markets, including different language communities, But broader comparison doesn't make the findings representative or turn mentions into percentages. Asynchronously participation can also help us reach people who are difficult to schedule, such as specialists, shift workers, or those in different time zones. But all of these are largely researcher benefits. I was curious to understand the experience and value for the person being interviewed. So, I conducted human moderation interviews, with 14 people who had previously taken part in AI moderated research. I heard that the clearest benefit was flexibility. As one participant explained, I can do it on my own time. I don't have to sit down specifically for a session. For people with a regular work, carrying responsibilities, or large time zone differences, that flexibility could be the difference between participating and not participating. So But flexibility doesn't necessarily mean a better conversation. All 14 participants spontaneously compared AI to human moderation. And no one described AI as the better conversation. Five, explicitly separated convenient access, from the interaction quality. So value was often not that AI provide us a better interview, but that it made the interview easier to access. Participants described another benefit. Reduced social pressure. As one participant explained, it's not the same pressure, having a human look at you while you're trying to formulate a response. For some, the absence of a visible human reaction felt more relaxed. And allowed greater directness. Even about mild embarrassment or frustration. So But AI doesn't simply make the interview less social. It removes social dynamic. So that can support candor, but it also removes a human moderator's ability to recognize the discomfort or distress and respond appropriately. So that's why I see the clearest fit in everyday, lowest state situations. Topics with reduced social pressure can support candor without creating a real need for human support. So The boundary isn't simply sensitive versus nonsensitive. Mild embarrassment may benefit from the absence of a visibly reacting person, but vulnerability or trauma is a different proposition. So the value is conditional, speed and breadth for researchers, flexibility, and sometimes less social pressure for the participant. So AI moderation earns its place when it suits the research question, is appropriate for the participant context, and can responsibly produce the kind of evidence we need. Choosing the use case is only the first design problem. The underlying model also shapes the conversation. Three characteristics are particularly relevant. And I saw all three of them across the sessions. First, limited unamazing context. In longer conversations, a moderator may lose track repeat a question, and forget where it was heading. Second, helpfulness and agreeableness. It may ask several questions, explain what the participant should interpret, lead them, and praise and answer. Third, plausibility over accuracy. The next question may fit the conversation, but introduce a detail or interpretation that the participant actually never supplied. Is that necessarily reasons to dismiss the method, but they are behaviors that we need to anticipate and design around. That is why much of the judgment moves upstream. A human moderator can notice drift, Aren't necessarily deeper. As conversations become more complex, a moderator is more likely to lose track, repeat itself, or pursue an unproductive thread. The second is the research goal. Each objective should be specific and narrow enough to explore properly in one session. Substantially different questions are best left for a separate study. And I also make sure to interrogate assumptions. If a given goal presumes a behavior or a problem, or that a particular feature is a is the answer to that problem, the moderator may carry that into the questions and follow ups. The third is context. The moderation needs enough background to understand the topic, but more isn't necessarily better. I keep only what helps to conduct the interview. Control of all of this does depend on the So Askable lets me structure objectives, adjust depth, and provide context. Although those controls evolve during the problem, program, which matters later. So principles apply across platforms, using them depends on controls and transparency we're given. The fourth lever is moderation guidance. How the moderator actually conducts the conversation. Researchers may control this directly, influence it indirectly, or leave it to the system. All of these principles should be familiar. The difference is making them explicit. Ask one open question at a time. Models often bundle several together. Focus on a concrete recent experiences, what someone did, what they might do, or what people generally think, Stay neutral. Don't praise, validate, explain, or advise even that's a great point may shape what follows. And don't assume human like warmth automatically creates trust. One participant I spoke to described AI trying to emulate a person. This trying to be my friend thing, it's not making me more comfortable. It's making me more distressful. Clarity and honest responsiveness may build more trust than scripted empathy. Finally, probe meaningful problems, workarounds, trade offs emotions, or us once or twice. Then move on. The participants I spoke to reinforced these principles, that it was easier for them to answer questions that were anchored in recent memories, and that the AI should treat, I've already answered this, as a stop signal, not as an invitation to rephrase. So None of this is new into your own practice. The challenge is also encoding it clearly enough for the system to follow. That leaves an important question. Did the moderator actually follow it? During the program, I was focused on reviewing what participants told us. So I went back and examined the moderator instead. I flipped the lens I treated the moderator as the subject of the review. What was it doing or failing to do across sessions? I reviewed over a thousand sessions from 10 research rounds. Across these four headliner dimensions shown here. Session conduct, probing behavior, and research coverage. Review was AI assisted. Accord code skill applied a defined framework to each transcript, flag behaviors such as repeated or leading questions, missed signals, broken closures, and left instructions, and then linked each flag to the evidence. The framework distinguished isolated slips from recurring problems, So a session won't be rated poor because of a minor imperfect in the conversation. Using AI to review AI obvious limitations. So I don't treat the scores as facts. The value is that I was able to review a large number of sessions, more consistent consistently and check each issue against the transcript. It also wasn't a controlled experiment. The studies varied, the platform changed, and my review method evolved as well. So I was looking for large patterns that showed up repeatedly rather than focusing too much on small differences. One of the clearest of these appeared when I compared the three rounds before Askabell introduced structured objectives. And depth controls with the seven rounds that followed. The most striking change was in research coverage. Or how successfully the moderator covered the study's objectives. The colors here show full partial, and insufficient coverage across each each round. Before structured objectives and depth controls, only 26 to 38% fully covered the study. The three lowest results. After the changes, full coverage increased to between 6991% in every round. Most of the early sessions were part partial, so the moderator explored much of the study, but most or compressed an objective. The pattern suggests that explicitly structuring each objective helped the moderator distribute as attention more reliably. Over the 10 rounds, question quality was consistently strong. And probing improved. Poor question quality ratings stayed between 08%, So question construction was really the main limitation. Now, In the three pre framework rounds, 10 to 22% of sessions receive poor probing ratings. After the changes, poor fell to between 08%, and the final three rounds were 70 97% to 98 good. So the moderator generally asks useful questions, and became more reliable at following up enough to meet each objective. The bigger remaining problem wasn't what it asked, it was whether it could run the session successfully. Session conduct was the most significant and persistent issue. Poor ratings range from two to 41%, with severe problems returning after apparently strong rounds. And the issue persisted even though its severity was intermittent. These were failures in running the conversation. So loops, internal state leaks, unacknowledged freezes, turn taking gaps, and broken endings. The structure framework appears to have improved coverage improving, but conduct was less responsive. Suggesting these failures involve session mechanics, as well as research design. So These rounds span changing versions of the platform, so it's not a claim that every session behaved this way. All that each issue reflects the current product. So, The final two rounds are encouraging. But the 10 round view shows why a clean recent result isn't enough to include that the risk had completely gone away. Two anonymized exchanges may help to make this category a bit more concrete. They are examples, so not representations of every session. Here, participant was waiting for the interview to continue. Instead, the moderator narrated, its time state and shut down instructions. They moved into a conventional closing statement. The session ended without asking the final research question. Was intermittent not typical, but it shows why conduct needs to be addressed separately. Useful quick questions don't guarantee a reliable interview. A second example shows the burden moving to the participant. The participant recognized a stall. And asked the moderator to continue. The moderator then repeated the participant's request back as though it were an answer. Then interpreted the complaint as their priorities, The participant had to correct it. Restate the request, and recover the conversation. The session still ended, so a completion measure might miss this. But the participant was doing the moderator's work. That raised a question the transcript couldn't answer. What did it cost participants to keep these interviews going? To understand that, I returned to the human moderated interviews with the 14 previous participants. The main finding wasn't dramatic frustration. It was how people seemingly absorb the failures. 10 described coping, repeating or rephrasing, prompting the moderator, shortening responses, or working around the expected behavior. Seven took responsibility for keeping the session moving. Recovering from dead ends, restarting, or checking that it had been captured, Yet 11 were willing to participate again. I Only two described overt frustration. And just one reached anger early termination, and permanent opt out. But visible escalation is only part of the story. In a live interview, an experienced researcher might notice their quieter adaptation and respond. Once removed from the session, we lose that opportunity Repair work can look like ordinary cooperation, while what participants decide not to say leaves no trace. So, One participant was careful about what he said, because he worried the AI would just disappear down an unproductive tangent. That cost is invisible. Session may look complete, the participant has edited his account to fit what he believes the moderator can handle. It wasn't simply a shorter or less detailed answer It was a genuine line of reasoning that never entered the transcript at all. An absence of visible frustration, may simply mean that correcting the moderator isn't worth the effort. So research coverage can appear successful, even though the participant has narrowed the evidence before it even reaches the transcript. This has changed how I assess the quality of an AI moderated research session. Coverage matters. But it isn't enough. Did the participant compensate for the moderation? Does the account reflect what they wanted to say? Rather than what the system can handle? A robust session needs all three of these. Research coverage, a reasonable participant experience, and evidence that we can trust. At the beginning, I borrowed the introduction that participants here in an AI moderated research session, there are no right or wrong answers. There may be no one correct way to use AI moderation. But there are better questions to ask of it. Not simply, how many interviews can we run? But what kind of evidence is this method structurally Not simply, did it complete the script? But did it cover the objectives, respect the participant, and produce trustworthy evidence. And not simply, what can AI do now, but what research judgment must surround it. Automation doesn't remove the researcher's responsibility. Moves their judgment upstream. Into the goal, context, boundaries, and guidance. It also creates a downstream responsibility. Review what happened, including how participants compensated. Some of this depends on researchers, some on what platforms expose and allow us to configure. A robust method requires us to test it, document failures, and be clear about the standards platforms must maintain. The interview may be automated, The judgment and responsibility for the quality of the research remains ours. Thank you. I've got, like, if, you're in confab and you can hear you've got a question, type it in and we'll be able to read it. Thank you. That sounds like a big program that you had to undertake. Mhmm. I I'm so just just something that I know. So one, I loved the rigor with with you reviewed the rigor of your method. So I thought that was lovely. It's something that I don't see at other conferences. It's one of the things that seems unique to design research. Research is the extent to which people in this community, in this audience, will go to that kind of effort to make sure that the methods we're using are actually delivering results that we can rely on. So appreciate that. Thank you. That was awesome. Does anyone have questions for Rebecca as we go. At the back there, and I'll repeat your question if that's okay. Go for it. Absolutely. What about the ethics of, having a computer ask So the question for the audience, at home is, has Rebecca looked at any of the looked into the question of the ethics of having a computer ask personal questions of a participant. I guess the so the program that I was running, the questions the the topics generally weren't delving into to kind of personal or private matters, it was more just general, without giving too much client information. Away, but general ecommerce. So I guess it wasn't touching too much into, private or personal topics. So that would be part of the process of in in my context, that would be part of the process of them signing up to be part of AskPaul's, panel. I guess they are choosing to be take part in sessions. They're aware it's a an AI, so it is a informed process. People have that option. Obviously, then there is also the aspect that, the the voices of the people who don't agree with that and aren't okay with that situation are I don't I don't Yeah. Anyone else have some questions? Don't have any online, but yes. Your research status. You mentioned about you were changing kind of the research objectives and going into more structural depth. So the question was around, whether or not Rebecca had to adjust or whether she did adjust her methods over time based on the way in which the tools were progressing. Yes. Yes. Definitely. There's definitely a bit of, I guess, experimentation on my side, changing what I input and then reviewing how that kind of affected the moderator's behavior. And then as well as the changes that askable made to their platform as well in parallel. Having those obvious impacts. So, I think that was one of the key learnings that I've had so far that, that is it's that working together and platforms providing that transparency so that you can understand what you're doing, what effect that's gonna have. Questionnaire. Yes? Track. Yeah. There is this response towards safety outside there. Yeah. Yeah. Sorry. And the question is around whether or not like, what safeguards are available to potential potentially close off lines of inquiry if we're getting into an area that is, potentially triggering or traumatic. But also, I I guess part of that is how well can we recognize that our participant is being traumatized. Would be a part of that? Yeah. Okay. Because these tools are pushed. Right? They are? These because these tools are pushed. Right? They are? These tools are pushed onto research. This is a good option. And it takes a human to kind of Yeah. I think that's that's currently missing that to provide that better experience, you probably need to be giving people more control to, you can you can reply and say, I've answered that or I'm not comfortable, and the moderator will respond, but I guess it's just having that that giving that

## Slides

### 00:00:34

# DEPTH at scale?

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:00:38

# DEPTH at scale?

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:01:00

# DEPTH at scale?
### Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:01:22

# DEPTH at scale?

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:01:48

## DEPTH at scale?
Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:02:14

## Different methods / different evidence

- How many - and among whom?
- What happened, in what order, and why?
- What reasons haven't we anticipated?
- What sits beneath stated experiences, across many people?

Adapted from Saeideh Bakhshi, "The fallacy of depth at scale", 2026

### 00:02:35

# Different methods / different evidence

## Survey
- How many - and among whom?
- ❌ Capture reasons you didn't ask about

## Depth interview
- What happened

### 00:02:37

# Different methods / different evidence

## Survey
- How many - and among whom?
- ❌ Capture reasons you didn't ask about

## Depth interview
- What happened

### 00:03:01

# Different methods / different evidence

## Survey
- How many - and among whom?
- **Missed:** Capture reasons you didn't ask about

## Depth interview
- What happened,

### 00:03:26

# Different methods / different evidence

## Survey
- How many - and among whom?
- ⓧ Capture reasons you didn't ask about

## Depth interview
- What happened, in what order, and why?
- ⓧ How common the experience is

## Open text
- What reasons haven't we anticipated?
- ⓧ What sits beneath the first answer

## AI moderation
- What sits beneath stated experiences, across many people?
- ⓧ Behaviour, the unsaid or prevalence

Adapted from Saeideh Bakhshi, "The fallacy of depth at scale", 2026

### 00:03:50

# Where AI moderation earns its place

1.  **Recurring or time-critical research**
    - Parallel sessions within a weekly cycle or product sprint

[Diagram with four numbered sections (1-4) arranged in a grid. Section 1 contains text, while sections 2, 3, and 4 are placeholders. Some sections feature abstract, glitch-like purple and blue patterns.]

### 00:04:14

# Where AI moderation earns its place

## 1 Recurring or time-critical research
- Parallel sessions within a weekly cycle or product sprint

## 2 Broader qualitative comparison
- How and why

### 00:04:32

# Where AI moderation earns its place
- **1 Recurring or time-critical research**
  - Parallel sessions within a weekly cycle or product sprint
- **2 Broader qualitative comparison**
  - How and why experiences differ across customer groups, markets and languages
- **3 Hard-to-schedule participants**
  - Specialists, shift workers or people across time zones
- **4**

[Four numbered sections, with abstract glitch-art images accompanying sections 1, 2, and 3. Section 4 is a blank black box.]

### 00:04:37

# Where AI moderation earns its place
- **1 Recurring or time-critical research**
  - Parallel sessions within a weekly cycle or product sprint
- **2 Broader qualitative comparison**
  - How and why experiences differ across customer groups, markets and languages
- **3 Hard-to-schedule participants**
  - Specialists, shift workers or people across time zones
- **4**

[Four numbered sections, with abstract glitch-art images accompanying sections 1, 2, and 3. Section 4 is a blank black box.]

### 00:05:01

> **I can do it in my own time.**
> I don't have to sit down specifically for a session.

Designing research for AI moderators | Rebecca Klee | Design Research 2

### 00:05:25

> **I can do it in my own time.**
> I don't have to sit down specifically for a session.

Designing research for AI moderators | Rebecca Klee | Design Research

### 00:05:32

> **I can do it in my own time.**
> I don't have to sit down specifically for a session.

Designing research for AI moderators | Rebecca Klee | Design Research

### 00:05:58

# DESIGN RESEARCH
## UX AUSTRALIA | WEB DIRECTIONS

> **I can do it in my own time.**
> I don't have to sit down specifically
> for a session

### 00:06:22

# DESIGN RESEARCH

> It's not the same pressure as having a human looking at you while you're trying to formulate a response.

Designing research for AI moderators | Rebecca Klee | Design Research 2026

[Line art illustration of a person's head with a thought bubble containing a small icon]

### 00:06:37

# Where AI moderation earns its place
- **1 Recurring or time-critical research**
  - Parallel sessions within a weekly cycle or product sprint
- **2 Broader qualitative comparison**
  - Exploring how and why experiences differ across customer groups
- **3 Hard-to-schedule participants**
  - Specialists, shift workers or people across time zones
- **4**

[Abstract, glitch-like images are shown in boxes 1, 2, and 3, and below box 3. Box 4 is a solid black rectangle.]

### 00:06:59

# Where AI moderation earns its place

1.  **Recurring or time-critical research**
    Parallel sessions within a weekly cycle or product sprint
2.  **Broader qualitative comparison**
    Exploring how and why experiences differ across customer groups
3.  **Hard-to-schedule participants**
    Specialists, shift workers or people across time zones
4.  **Everyday, lower-stakes topics**
    Situations where reduced social pressure may support candour

[Four abstract, glitch-like images are placed next to each numbered point]

### 00:07:21

# Where AI moderation earns its place

1.  **Recurring or time-critical research**
    Parallel sessions within a weekly cycle or product sprint
2.  **Broader qualitative comparison**
    Exploring how and why experiences differ across customer groups
3.  **Hard-to-schedule participants**
    Specialists, shift workers or people across time zones
4.  **Everyday, lower-stakes topics**
    Situations where reduced social pressure may support candour

[Abstract, glitch-like images are used as visual elements within some of the numbered sections.]

### 00:07:36

# Where AI moderation earns its place

1.  **Recurring or time-critical research**
    Parallel sessions within a weekly cycle or product sprint
2.  **Broader qualitative comparison**

### 00:07:43

UX AUSTRALIA | WEB DIRECTIONS
# DESIGN RESEARCH

# Where AI moderation earns its place

1.  **Recurring or time-critical research**
    Parallel sessions within a weekly cycle

### 00:08:08

# How AI model characteristics show up

- **Limited context**
  Forgets earlier answers, repeats questions or loses the thread
- **Helpful and agreeable**
  Over-explains, asks too much, leads or praises participants
- **Plausibility over accuracy**
  Introduces details or interpretations the participant never supplied

Designing research for AI moderators | Rebecca Klee | Design Research 2026

[Icon of two horizontal arrows pointing left and right, inside a square]
[Icon of a heart with an exclamation mark]
[Icon of a stylized asterisk or starburst]

### 00:08:31

# How AI model characteristics show up

- **Limited context**
  Forgets earlier answers, repeats questions or loses the thread
- **Helpful and agreeable**
  Over-explains, asks too much, leads or praises participants
- **Plausibility over accuracy**
  Introduces details or interpretations the participant never supplied

*Designing research for AI moderators | Rebecca Klee | Design Research 2026*

[Three circular icons: one with a square and two horizontal arrows, one with a heart and an exclamation mark, and one with a stylized asterisk.]

### 00:08:32

# How AI model characteristics show up

- **Limited context**
  Forgets earlier answers, repeats questions or loses the thread
- **Helpful and agreeable**
  Over-explains, asks too much, leads or praises participants
- **Plausibility over accuracy**
  Introduces details or interpretations the participant never supplied

Designing research for AI moderators | Rebecca Klee | Design Research 2026

[Three circular icons: one with horizontal arrows, one with a heart and exclamation mark, and one with a starburst/asterisk shape.]

### 00:08:55

# Move the judgement upstream
Four levers shape the session before it begins

- **Time and depth**
- **Context**
- **Research goal**
- **Moderation guidance**

Designing research for AI moderators | Rebecca Klee | Design Research 2026

[Four rounded rectangles, each containing an icon and text: a stopwatch for "Time and depth", a list for "Context", a target for "Research goal", and a question mark in a speech bubble for "Moderation guidance".]

### 00:09:22

# Set the boundaries

## Time and depth
Limit the session and the time spent on each goal.

## Research goal
Keep it focused, specific and free from assumptions.

## Context
Provide only the what's needed to inform the interview.

[Icon of a stopwatch]
[Icon of a target with a magnifying glass]
[Icon of a document with bullet points]

### 00:09:48

# Set the boundaries

- **Time and depth**
  Limit the session and the time spent on each goal.
- **Research goal**
  Keep it focused, specific and free from assumptions.
- **Context**
  Provide only the what's needed to inform the interview.

[Icon of a stopwatch]
[Icon of a target with a magnifying glass]
[Icon of a document with bullet points]

### 00:10:11

## Set the boundaries

-   **Time and depth**
    Limit the session and the time spent on each goal.
-   **Research goal**
    Keep it focused, specific and free from assumptions.
-   **Context**
    Provide only the what's needed to inform the interview.

[Icon of a stopwatch]
[Icon of a target with a magnifying glass]
[Icon of a document with a bulleted list]

### 00:10:33

# Set the boundaries

**Time and depth**
Limit the session and the time spent on each goal.

**Research goal**
Keep it focused, specific and free from assumptions.

**Context

### 00:10:57

# Make interviewing discipline explicit

## Do...

## Don't...

### 00:11:20

# Make interviewing discipline explicit

## Do...
- **Ask one open question at a time**

## Don't...
- Ask compound or leading questions

### 00:11:33

# Make interviewing discipline explicit

## Do...
- Ask one open question at a time
- Focus on concrete, recent experience
- Stay neutral

## Don't...
- Ask compound or leading questions
- Ask for hypotheticals or generalisations
- Praise, validate or explain

### 00:11:56

# UX AUSTRALIA | WEB DIRECTIONS
## DESIGN RESEARCH
### Make interviewing discipline explicit

#### Do...
- Ask one open question at a time
- Focus on concrete, recent experience
- Stay neutral

#### Don't...
- Ask compound or leading questions
- Ask for hypotheticals or generalisations
- Praise, validate or explain

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:12:19

# Make interviewing discipline explicit

## Do...
- Ask one open question at a time
- Focus on concrete, recent experience
- Stay neutral
- Probe useful signals once or twice

## Don't...
- Ask compound or leading questions
- Ask for hypotheticals or generalisations
- Praise, validate or explain
- Follow every thread indefinitely

### 00:12:35

# Make interviewing discipline explicit

## Do...
- **Ask one open question at a time**
- **Focus on concrete, recent experience**
- **Stay neutral**
- **Probe useful signals once or twice**

## Don't...
- Ask compound or leading questions
- Ask for hypotheticals or generalisations
- Praise, validate or explain
- Follow every thread indefinitely

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:12:41

# Make interviewing discipline explicit

## Do...
- **Ask one open question at a time**
- **Focus on concrete, recent experience**
- **Stay neutral**
- **Probe useful signals once or twice**

## Don't...
- Ask compound or leading questions
- Ask for hypotheticals or generalisations
- Praise, validate or explain
- Follow every thread indefinitely

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:13:02

# 1000+ sessions

AI-assisted review across:
- Session conduct
- Question quality
- Probing behaviour
- Research coverage

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:13:25

# 1000+ sessions

## AI-assisted review across:
- Session conduct
- Question quality
- Probing behaviour
- Research coverage

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:13:31

# 1000+ sessions

## AI-assisted review across:
- Session conduct
- Question quality
- Probing behaviour
- Research coverage

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:13:56

# 1000+ sessions

AI-assisted review across:
- Session conduct
- Question quality
- Probing behaviour
- Research coverage

Designing research for AI moderators | Rebecca K

### 00:14:20

# 1000+ sessions

AI-assisted review across:
- Session conduct
- Question quality
- Probing behaviour
- Research coverage

Designing research for AI moderators | Rebecca K

### 00:14:34

### Designing research for AI moderators
# 1000+ sessions
AI-assisted review across:
- Session conduct
- Question quality
- Probing behaviour
- Research coverage

### 00:14:58

# DESIGN RESEARCH
## UX AUSTRALIA | WEB DIRECTIONS

### Research coverage
Objectives fully covered — Full / Partial / Insufficient scale

- Full - all objectives covered
- Partial -

### 00:15:22

# DESIGN RESEARCH

## Research coverage

-   **Full**: all objectives covered
-   **Partial**: some objectives only partly covered
-   **Insufficient**: objectives left largely uncovered

Objectives

### 00:15:48

# Designing research for AI moderators
[Two stacked bar charts showing quality metrics, one for 'Question quality' and one for 'Probing behaviour', with categories 'Good', 'Fair', and 'Poor

### 00:16:12

# UX AUSTRALIA | WEB DIRECTIONS
## DESIGN RESEARCH

- Good: ≥ 75% turns clean — no recurring problem
- Fair: a problem recurs (~10%+ turns)
-

### 00:16:35

### Designing research for AI moderators
#### Session conduct
- **Good**: ≥75% turns clean - no recurring problem
- **Fair**: a problem recurs (~10%+ turns)

### 00:17:02

# Session conduct
Script loops, internal-state leaks, unacknowledged freezes

**Legend:**
- Good: >75% turns clean - no recurring problem
- Fair: a problem recurs (~10%+ turns)
- Poor: severe recurrence (~25%+ turns)

INTERMITTENT

Designing research for AI moderators | Rebecca Klee | Design Research 2026

[Stacked bar chart showing session conduct over 10 sessions, with each session's performance categorized into good, fair, and poor recurrence rates, displayed as percentages.]

### 00:17:25

# UX AUSTRALIA | WEB DIRECTIONS
## DESIGN RESEARCH

### Session conduct
- **Good**: >75% turns clean - no recurring problem
- **Fair**: a problem recurs

### 00:17:32

# UX AUSTRALIA | WEB DIRECTIONS
## DESIGN RESEARCH

### Session conduct
- **Good**: >75% turns clean - no recurring problem
- **Fair**: a problem recurs

### 00:17:56

# Session conduct
Script loops, internal-state leaks, unacknowledged freezes

**Legend:**
- Good: >=75% turns clean - no recurring problem
- Fair: a

### 00:18:19

# Designing research for AI moderators

**Participant:** Is there anything else?
**AI Moderator:** Internal time status. Time status. L-A-T-E-S-T-

### 00:18:34

### Designing research for AI moderators
[Screenshot of a web browser displaying a Google Slides presentation with a black slide and a footer]

### 00:18:42

### Designing research for AI moderators
[Screenshot of a web browser displaying a Google Slides presentation with a black slide and a footer]

### 00:19:06

**AI Moderator:** You're looking for genuine value that aligns with your actual needs, not just extras for the sake of it. Got it. That's really helpful.
**[1

### 00:19:31

## The participant may be doing the repair work.

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:19:33

## The participant may be doing the repair work.

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:19:58

The participant may be doing the repair work.

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:20:23

## The participant may be doing the repair work.

### 00:20:31

## The participant may be doing the repair work.

### 00:20:56

# Designing research for AI moderators

> "**I'm very careful about what I say** because I don't want to accidentally set it off ... down some tangent."

[Illustration of

### 00:21:21

UX AUSTRALIA | WEB DIRECTIONS
# DESIGN RESEARCH

> " **I'm very careful about what I say** because I don't want to accidentally set it off ... down some tangent."

### Designing research for AI moderators
Rebecca Klee | Design Research 2026

[Illustration of a person's head with a beard and headphones]

### 00:21:34

> **I'm very careful about what I say** because I don't want to accidentally set it off ... down some tangent.

Designing research for AI moderators | Rebecca Klee | Design Research 2026
[Illustration of a person's head with a beard and headphones]

### 00:21:56

### 1 Research coverage
Did we explore what we intended to learn?

### 2 Participant experience
Was the participant's time, effort and contribution respected?

### 3

### 00:22:20

# What this asks of us

### 00:22:33

## What this asks of us

### 00:22:39

## What this asks of us

### 00:23:00

# What this asks of us
- **Choose** the method for its structural fit.

Designing research for AI moderators | Rebecca Klee | Design Research 2026

### 00:23:24

# What this asks of us

- **Choose** the method for its structural fit.
- **Move** research judgement upstream.
- **Evaluate** coverage, participant experience and evidence quality.

### 00:23:50

# DESIGN RESEARCH

> The interview may be automated.
> The judgement isn't.

Designing research for AI moderators | Rebecca Klee | Design Research 2026

[Slide

### 00:24:14

# Thank you

I acknowledge the traditional owners of the land from which I present, the Traditional Custodians of this land, and pay my respects to their Elders past and present.
I would like to thank you to John, Marcus, Sam, Holly and the team for their dedication for hosting this conference.
This presentation was brought to life using photography by Igor Kisselev and Logan Nao on **Unsplash**. Alongside icons from the **Noun Project**.

Designing, research for AI moderation | Rebecca Rice | Design Research 2024
@rebeccarice_design

[QR code with text "Presentation slides" below it]

### 00:25:02

[Photograph of a dense green forest or garden with a small, light-colored sign visible in the lower-middle left.]

### 00:28:33

### NSW TEACHERS FEDERATION

### 00:29:19

[The slide content is illegible due to blur and distance.]

### 00:29:32

[Photograph of a group of people, possibly students or colleagues, in a bright, modern setting.]

### 00:29:57

[Extremely blurry and cut-off graphic on a white slide, content unidentifiable]
