This article presents the methodology we used in our study of how professionals use AI text-generated tools like ChatGPT.

Research Method

The objective of our study was to uncover user behaviors and usability issues faced by working professionals interacting with AI text-generation tools like ChatGPT.

We conducted 8 90-minute moderated qualitative usability tests. We chose this method so we could naturally observe user-behavior patterns in a relatively short amount of time. We considered other methods like contextual inquiry and interviews but they do not place as much emphasis on uncovering usability issues.

We used a think-aloud protocol with participants. We asked probing questions to understand the intent behind participants’ actions and uncover underlying mental models.

Participant Profile

We strategically set broad recruiting parameters, looking for y working professionals with experience using text-based AI tools for work.

We recruited a total of 8 participants:  6 women and 2 men ranging between 26 and 55 years in age. Their professions included project management, marketing, design, writing, consulting, and research. All of them reported in the screener survey that their most-used generative AI tool was ChatGPT and they were using it a few times a week or more. We asked participants to bring to the test a few tasks of their own that they had not attempted so far with ChatGPT. Participants brought tasks that they wanted to complete as part of their jobs like creating text for a brochure for their company or creating an itinerary for an upcoming business trip.

Session Structure

The usability-testing sessions were conducted virtually over Zoom. Each session lasted 90 minutes and was divided into 3 parts:

  1. Initial interview (10–20 minutes). The first 10-20 minutes were a discussion about what participants did day to day as part of their work. Understanding this helped gauge which of the tasks we had prepared for the participants would be the most relevant.
  2. Participant’s task (10–20 minutes). The second part of the session was having participants attempt tasks they brought to the session. This part helped participants get comfortable with the setup of the study and the tool. Participants brought tasks like: making a travel itinerary for a trip, creating text for a brochure based on some slides from their workplace, <more examples>.
  3. Predefined, participant-specific tasks. The third and longest part of the session was dedicated to having participants perform predefined tasks using ChatGPT. These predefined tasks were custom-created by us, to consider the participant’s background gathered via the screener. The participants were provided an option to use their own ChatGPT account or use a test account we had set up for them.

Creating Relevant Tasks

Formulating custom tasks for each participant was challenging, as we had participants from a broad range of professions and, thus, a wide range of day-to-day tasks and goals. We aimed to create tasks that would be relevant and realistic for each participant’s work.

We approached this challenge by:

  1. Asking participants to bring a task of their own
  2. Using the screener responses to create tasks
  3. Getting ChatGPT’s help

1. Asking Participants to Bring Their Own

As part of the recruitment process, we asked the participants to bring a task of their own that they had not previously attempted with ChatGPT.

Using participant-provided tasks can introduce bias within the study as the participants may think of ahead of time and even try them by themselves. To prevent this possibility, we asking them to not attempt these tasks beforehand and we focused mostly on the tasks we created to draw our conclusions.

However, starting with their own tasks helped participants feel comfortable and gave us valuable insights on how they used ChatGPT. Based on this information, we were able to guide the session towards tasks that fit their context. After the first task, we selected relevant tasks from our list of custom, predefined tasks. 

2. Using the Screener Responses to Design Custom Tasks

Our screener asked potential participants questions about their day-to-day work and responsibilities, including  what kind of text-based artifacts they worked on in their jobs and what they had previously used generative-AI tools for. Based on this information, we created realistic, relevant tasks for each participant.

For example, here’s a participant’s response to our screener question:

What is your job title?

Marketing Director

What are some of the work-related tasks you’ve used generative AI tools for?

Social media posts, website legality text, website copy, blog posts

Given this context, we created a scenario around a fictional startup that had a marketing need and asked the participant to work on marketing deliverables:

Scenario:

A startup wants to launch a wellness app called ‘Mynd & Bodi’ with a focus on yoga, mindfulness, and therapy. It offers many online programs that customers can follow at home and services to get guidance from wellness coaches. The company’s priority is to get signups for the new app in an already competitive space where many wellness apps like Headspace are already popular. The target audience for the app consists primarily of adults over 40.

Tasks:

  • Find some ideas around how you’d approach this as a marketing challenge.
  • Create a marketing strategy for this new product launch.
  • Make some marketing materials that would promote the app on social media. You are free to choose the format/medium.
  • You’ve gotten feedback from the company that the materials you’ve made should convey that the services they offer are research-backed and therefore should be trusted. Please make the materials you’ve made convey the feeling of trustworthiness for customers.
  • The company wants to launch the app in Spanish as well. Please change the marketing materials to include Spanish.

3. Getting ChatGPT’s Help

We also used ChatGPT to help us create scenarios and tasks for specific jobs. We selected the best ideas and refined them to arrive at our final set of scenarios and tasks.

A screenshot of a ChatGPT prompt to create a realistic project scenario.
An example ChatGPT prompt we used to generate a scenario and tasks for a project manager

Covering All Task Types and Fidelity Levels

To ensure the validity of our findings, we wanted to cover as many use cases of AI text-generation tools as  possible. As part of our secondary research, we examined prompt datasets online and lists of tasks that participants performed with AI chatbots in  another, earlier diary study documented in our article  Information Foraging with Generative AI: A Study of 3 Chatbots. Based on the different tasks we observed, we came up with a list of 5 possible use cases that any text generation task would fall under:

  1. Coming up with ideas and inspiration
  2. Creating text content for a deliverable (business document, social media post, assignment, article, marketing copy)
  3. Learning by asking the AI to explain a topic
  4. Retrieving and compiling factual information about a topic
  5. Asking the chatbot to modify or interpret a piece of text

This general list ensured that we covered a range of interesting tasks for each user and saved us time.

We also made sure that all these use cases were represented in the tasks we created for each participant. It is important to note that a task could cover more than one of these use cases if it was complex enough to involve multiple steps.

Use case

Task examples in the context of marketing

Coming up with ideas and inspiration

Find some ideas around how you’d approach this as a marketing challenge.

Creating text content for a deliverable (business document, social media post, assignment, article, marketing copy)

Make some marketing materials that would promote the app on social media. You are free to choose the format/medium.

Learning by asking the AI to explain a topic

Is there a particular topic you’ve been aiming to learn more about as part of your work? Any open questions you have? Find out more about it using AI.

Retrieving and compiling factual information about a topic

Create a blog post for beginner marketing professionals about copyright, infringement, and website legality so they can learn best practices.

Providing the chatbot text and asking it to modify, interpret or deduce from it

The company wants to launch the app in Spanish as well. Please change the marketing materials to include Spanish.

It was important to ensure that the deliverables we ask participants to work on spanned all levels of fidelity. (Approaches to generating text could be different based on what type of output is being targeted.) We defined 3 levels of fidelity, which we aimed to cover for each participant:

  • Low-fidelity outputs: The content takes priority, and the format and writing style are unimportant (ideas, inspiration, etc.).
  • Mid-fidelity outputs: Formatting is somewhat important, but the writing style does not have to be polished (article outlines, marketing plans, etc.).
  • High-fidelity outputs: These require appropriate formatting and high-quality writing style (emails, social media posts, articles, etc.).

Taking the above example of the tasks for the marketing director, we can see all three levels:

Fidelity level

Task examples in context of marketing

Low

Find some ideas around how you’d approach this as a marketing challenge.

Mid

Create a marketing strategy plan for this new product launch.

High

Make some marketing materials that would promote the app on social media. You are free to choose the format/medium.

Conclusion

Our approach to task creation proved to be a successful strategy, as all participants responded well to the tasks we asked them to attempt. During the session, we gave participants the opportunity to provide criticism of any task that they considered unrealistic, but we did not receive any such complaints. We were able to cover all types of tasks while maintaining relevance to the participants’ contexts. Despite having a small sample size of 8, the study surfaced numerous user behaviors and usability issues, which have been covered in our main articles.