02 — ACE Lab, USC · UIST ’25

Beyond the Page

Enriching academic paper reading with peer discussions from social media — without the cognitive overload.

My Role
HCI Researcher & Co-author
Timeline
2024 — 2025
Team
5 Researchers, ACE Lab USC
Venue
UIST ‘25, Busan
Read the paper → UIST ’25
§ Overview

What is SURF?

SURF — Social Understanding of Research Findings — is a novel paper reading interface that enriches academic literature with peer discussions pulled from social media platforms like X (formerly Twitter). The project was conducted at the Adaptive Computing Experiences Lab at the University of Southern California and published in UIST 2025.

As a co-author and HCI researcher on a five-person team, I contributed to every stage of the research lifecycle: from running the formative study and synthesizing design goals, to informing the interface design and conducting comparative usability testing. This was my first published research paper at an ACM top-tier venue.

The core question

Researchers actively share informal, accessible breakdowns of papers on social media — tweetorials, Q&As, critiques, perspective threads. These conversations are often more illuminating than the paper itself, providing extra context and asking critical questions. But they exist completely separately from the act of reading. SURF asks: what if they didn’t have to?

The Problem

A gap between paper and people

Academic papers are written to uphold rigorous standards of accuracy and reproducibility. This makes them dense, formal, and difficult to parse without significant prior knowledge. At the same time, the sheer volume of papers being published keeps growing, pushing researchers to skim faster and retain less.

Meanwhile, on platforms like X and BlueSky, vibrant communities of researchers share informal, accessible breakdowns of those same papers — using examples, personal reflections, and multimedia content to lower the barrier. These conversations are rich with insight. But they go almost entirely unused during the actual act of reading.

Four tensions we identified

01
The context-switch problem

Switching between a dense PDF and a social feed is disruptive. Researchers lose focus, lose thread, and often don't bother.

02
The noise problem

Social media is full of hype and shallow commentary. Finding the correct signal requires effort most readers won't spend mid-read.

03
The credibility problem

Without context, informal posts feel unreliable — even when they're as useful as the paper's own prose.

04
The thread problem

Academic discussions branch into nested subthreads. Valuable exchanges get buried, and following them demands the same attention reading already requires.

Research

Formative study

Before designing SURF, we ran a formative study with 8 researchers — 7 PhD students and 1 master’s student, across HCI, NLP, and ML fields. We built a functional technology probe: a basic interface that placed categorized social media threads alongside a research paper, with color-coded links between posts and the paragraphs they referenced.

Each session lasted 90 minutes — 35 minutes of reading with the probe, followed by a semi-structured interview. We observed how participants moved between formats, what they valued, and what frustrated them.

What we found — the benefits

Deeper comprehension

All 8 participants said social media helped them grasp the full picture of a paper — including its caveats and limitations that formal writing tends to bury.

Literature discovery

7 of 8 noted that social media discussions surfaced related papers, talk videos, and direct paths to authors they'd have otherwise never found.

Critical depth

3 participants found that threads surfaced implementation details and caught false causal claims the paper itself obscured.

What we found — the challenges

C1
Finding credible discussions

5 of 8 participants worried about hype, misinformation, and influential bad takes — particularly from what one participant called “tech bros” and “research influencers.”

C2
Following long threads

Hierarchical, branching conversations were hard to follow. Insightful exchanges were buried in nested replies that participants skipped entirely — “because they are very scattered.”

C3
Information overload

Too many discussions at once was distracting and created cognitive friction. Some felt seeing others' opinions early biased their own understanding before they could form it.

“Some people are asking technical details, like the actual implementation or model choice in their experiment — that would be something you get from the comments more than from the paper itself.”
— Participant T3, Formative Study
Design Goals

Five principles to guide the design

The formative study surfaced a central tension: participants saw real value in social media discussions, but the cognitive cost of engaging with them was consistently too high. We distilled five design goals to resolve it.

DG1
Support fluid movement between formats

Surface only the most relevant discussions and provide contextual links in the PDF — so readers can move between paper and social media without breaking focus or building a mental map from scratch.

DG2
Accommodate diverse reading styles

Some readers want to anchor on discussions first; others want to form their own view before seeing peer opinions. Allow readers decide when and how social context enters the experience.

DG3
Structure discussions for readability

Hierarchical threads are great for debate, but terrible for one-time comprehension. Flatten branching conversations into linear, skimmable storylines that readers can parse in step with the primary text.

DG4
Build trust in informal sources

Highlight valuable discussions, surface author credentials, and cut through noise. Make social media discourse feel as reliable as it often actually is.

DG5
Minimize visual distraction and cognitive load

Use progressive disclosure and low-saturation signifiers to surface contextual information without overwhelming the primary reading experience.

The Interface

Four features that do the work

SURF places the paper on the left and social media threads in a right-hand panel. Discussions are linked to specific paper sections via small icons, and readers can navigate between the two formats by clicking those links. The system uses consistent color coding so visual connections between threads and sections are intuitive at a glance.

01 — Faceted Linkage

Small icons placed next to section titles represent different discussion types: Overview, Q&A, Critique, Perspective, Related Work, and more. Clicking one filters the right panel to show only that type of thread for that section. Clicking the tag above a tweet scrolls the paper to the corresponding paragraph — bidirectional navigation that supports any reading strategy.

02 — In-situ Discussion Summaries

Hovering over a linkage icon reveals a tooltip with a concise narrative summary of key discussions around that section. Readers can preview whether a thread is worth diving into before committing their attention. Clicking through expands the full accordion and surfaces every linked thread.

03 — Scaffolded Navigation via Overview Threads

Tweetorials — threaded posts that walk through a paper step by step — are surfaced as structured overviews. Each step is linked to its corresponding paper section, giving readers a casual-language alternative to the abstract. In the usability study, nine participants preferred using this to the abstract for initial orientation.

04 — Focus Mode

A toggle between Focus Mode — highest-quality discussions only — and Social Mode, a less filtered view. SURF assigns each post a quality score based on how much it deepens understanding, broadens perspective, or makes dense content accessible. Hovering over user avatars surfaces their X profile, including institution and bio, so credibility is visible at a glance without leaving the interface.

§ Outcomes

What we learned from 18 researchers

We ran a within-subject comparative usability study with 18 researchers, comparing SURF against a plain PDF reader. Participants completed two reading sessions on NLP papers they hadn’t seen before, then wrote a mini-review identifying key strengths, weaknesses, and critical questions. Two expert raters evaluated the reviews blind to condition.

↑ 0.56
Soundness (p = 0.036)
↑ 0.48
Cognitive depth (p = 0.045)
↑ 0.53
Insightfulness (p = 0.046)
↓ 0.74
Mental demand (p = 0.036)
↓ 0.83
Frustration (p = 0.004)
17/18
Said SURF improved understanding

How people explored

All 18 participants actively switched between paper and social media discussions in SURF, averaging 12 context switches per session. Three distinct exploration styles emerged: discussion-first (scan discussion threads to orient before reading), paper-first (read fully, then turn to discussions), and interleaved (move fluidly between both throughout).

15 participants used the faceted linkage feature to navigate between linked tweets and paper sections. 16 used the in-situ summary feature an average of 6 times per session. 7 toggled off Focus Mode briefly but returned to it — realizing it filtered exactly the content they didn’t want to see.

“Whenever you start a paper and the topic is new for you, you might have silly questions... those naive questions have already been asked and answered on social media. It would be beneficial to have both expert questions and naive questions together.”
— Participant P6, Usability Study

Beyond reading

9 participants said SURF made them more motivated to engage in academic social media themselves. 11 said it helped them discover related literature faster than conventional search. 4 said they’d be more likely to DM authors’ social media accounts directly linked through the interface than hunt for an email address.

§ Reflection

What I took away

Cultivate trust, or lose the user

Informal sources carry little default credibility. Without signals of who’s speaking and why they might be worth listening to, users opt out entirely — regardless of how useful the content is. Every design decision in SURF that surfaced author credentials, institutional affiliation, or quality scores was doing trust-building work. It mattered.

Flatten the thread

Conversation threads are built for debate, not for just-in-time comprehension. When integrating threaded content into a reading experience, the system has to distill branching chatter into linear, skimmable storylines. This generalizes: any product fusing information sources of sharply different density faces the same problem.

What I’d do differently

The study used a 45-minute reading window, which doesn’t reflect real-world deep reading over days. Our participant pool skewed toward technically literate researchers. And SURF was tested primarily on NLP papers, where social media coverage is much denser than most academic fields. A longitudinal field study might tell a different and probably richer story.

What’s next

The bigger question — whether casual social commentary could seed a new paradigm of informal peer review — is still open. We’re interested in expanding SURF to BlueSky, integrating with open review platforms like OpenReview, and studying its effects across research communities where social media engagement looks very different.