Research & Behavioral Systems

AI Misattribution

In Study One, 123 participants evaluated an identical transcript under three stated-origin conditions. Only the label changed.

Academic Research
Diagram of one transcript evaluated under three different origin labels

Context

The work in context.

The study asked whether expectations about who produced a conversation can influence social judgments even when the words are unchanged.

Case study summary

What did Juan build?

An online experiment testing whether labels about AI involvement change judgments of the same interaction.

This case study was last reviewed on . Project details and limitations are identified throughout the page.

Role, audience and objectives

Responsibility carried through the system.

Juan's role

  • Research question and literature review
  • Experimental materials and questionnaire
  • Online study administration
  • Qualtrics implementation
  • IBM SPSS analysis
  • Interpretation and academic writing

Audience

  • Researchers studying human–AI interaction
  • Designers and teams communicating AI involvement

Objectives

  • Hold the transcript constant
  • Vary only the stated origin of the interaction
  • Measure perceived naturalness, humanness, and thoughtfulness

Behavioral lens

  • Context and expectations can affect evaluation
  • Measured outcomes should remain tied to the study conditions

Study One

One hundred twenty-three participants were assigned equally across three conditions: Human–Human, Human–AI, and AI–AI. Everyone read the same approximately 446-word transcript online. Only the stated origin changed.

Naturalness differed across conditions, F(2, 120) = 8.05, p < .001. Participants rated the Human–Human version more natural than either version described as involving AI. The Human–AI and AI–AI conditions did not differ significantly.

Study Two

The extension crossed responder identity with warm or cold tone. It asked whether tone changed evaluation in the same way for a response labeled human and a response labeled AI. The clearest pattern was that warmth improved ratings for the human-labeled response; it did not produce the same change for the AI-labeled response.

Interpretation

The findings support a careful conclusion: beliefs about origin can shape social evaluation in this text-based context. They do not show that people always reject AI, and they do not imply that labels control judgment.

Limitations

The studies used convenience university samples and a single academic interaction. Future work could test different relationships, settings, voices, avatars, and levels of perceived autonomy.

Information architecture and visual direction

Important structural decisions.

Information architecture

  • Consent and online questionnaire
  • Randomized label condition
  • Identical transcript
  • Six-point evaluation scales
  • SPSS analysis

Design decisions

  • Used an identical underlying transcript to isolate the label manipulation
  • Included Human–Human, Human–AI, and AI–AI conditions
  • Extended the work with a warm-versus-cold tone study

Technical implementation

Infrastructure, responsive behavior, and access.

Email, forms, payments and platform

  • Qualtrics questionnaire and participant materials
  • Online study administration
  • IBM SPSS data analysis

Accessibility and responsive design

  • Plain-language public summary
  • Labels written out before abbreviations
  • Findings described without implying determinism

Available evidence

What can—and cannot—be claimed.

Available outcomes

  • Study One found higher naturalness ratings in the Human–Human condition than in either condition labeled as involving AI

Limitations

  • Convenience university sample
  • Single text-based academic interaction
  • Results should not be generalized to every AI interaction

Reflection

What the project clarified.

The work reinforced that people do not evaluate language in a vacuum. The surrounding explanation can change how the same exchange is experienced.