Noctua Proxima

Scribblings

6 Questions to Ask When You See "Research-Backed"

A lot of claims today come labeled "research-backed" — grounded in data, drawn from real usage, presented with the confidence of a finding. That label alone is worth pausing on: is there actually research behind it, and if so, is it any good? Most readers don't have a research methods background, and don't need one to get to the bottom of these questions — you just need to know what to look for...

As a case study, consider Anthropic's piece Getting good at Claude: A research-backed curriculum, which describes patterns from "over 50,000 conversations across the 11 behavioral AI fluency  indicators" to say something about how people may become more fluent in using AI (Claude in particular). This piece is worth practicing on not because of any glaring errors or wild claims, but because it's a pretty typical example of research-flavored writing published by a company with a stake in the conclusion. Nobody's trying to swindle you here — they want to share information that leaves you with a favorable view of their product, and that's a perfectly valid motivation — companies need to stay in business. Where the piece falls short, it falls short in ordinary ways: methods go unreported, the kind of omission that's routine in a company blog post but would raise eyebrows in a peer-reviewed journal. Nothing extreme. And yet — mundane as it is — how should a reader judge a claim like this?

What follows are six questions for working through a piece like this — applied here to this specific article, but reusable on whatever you read next.

1. Who wrote this, and what do they gain if you believe it?

Before evaluating any claim, it helps to know who's making it and what they have at stake. Knowing the author lets you check their track record and expertise. Knowing who benefits if you believe the claim tells you how much scrutiny to bring.

With this article, however, there's no author listed, no contact information, and so no one to ask a follow-up question, and no individual putting their professional judgment behind the conclusions. What is clear is who published it: Anthropic, the company that makes Claude, in a piece about how to get more out of Claude. That doesn't necessarily mean the claims are biased, but it does mean Anthropic has something to gain from publishing it. An independent academic study or a journalist covering the same ground doesn't have the same stake in how the article is received. Worth reading with that difference in mind, rather than assuming the piece was written purely to share knowledge.

2. Can you state the actual claim in one plain sentence?

Before judging whether a claim is well-supported, you have to be able to say what the claim is. A good test: try compressing the article's central finding into one plain sentence, stripped of its specific vocabulary (jargon), and see what's left.

Here is my guess at the claim: people get good at Claude ("fluency") by practicing certain behaviors ("signature moves"), and which behaviors matter most differs somewhat by product (i.e., Chat, Cowork, etc.). If that's the claim, does it actually give you, the reader, any new information you can use? Practice helps people improve at most things, and different tools offer different functionality, so naturally users engage in different behaviors to fit that functionality. The specific words doing the work in the article — "signature move," "rewards," "curriculum," "fluency indicators" — mostly disappear once the claim is restated in plain language. It may be worth asking whether they're adding real precision with those terms, or possibly making the claim sound more substantial than it is.

Good research writing states the claim it's testing — the question it sets out to answer — up front, followed by a sentence or two describing the answer it found. From there, it goes into more detail about how it arrived at that answer: a description of the data (which this article provides), a description of the methods used to analyze it (which this article doesn't), and an interpretation of the findings in the context of other research (which this article gestures at, by linking to related product tutorials, but not by engaging with comparable independent research on topics around how people typically learn). Given those gaps, I found it genuinely difficult to state the central claim in one plain sentence.

As the reader, it's natural to default to assuming the fault is yours. But the greater burden sits with the author: writing clearly for your audience is the author's job, not a test of the reader's inference skills. The audience for this piece is broad — anyone who uses Claude or is curious about it — and writing for that range of readers means the concepts should be plain-spoken and the overall claim explicit, not something readers need several passes to pin down. Because the claim here was difficult for me to tease out, I read the rest of the piece more critically, trying to pin down exactly what the author was trying to show, how they went about it, and what they concluded. Part of what makes this piece hard to compress is that it isn't entirely clear what type of claim is being made. Is it making a causal claim or simply describing two things that tend to occur together — which is worth checking directly.

3. What kind of claim is being made? Is this a causal claim or a correlation, and does the piece tell you which?

This article makes a claim about how things relate to each other. Relationships like this usually come in two kinds:

Causal

  • X produces Y —It rained today, so the ground is wet.

Correlational

  • X and Y co-occur without one causing the other — People carry umbrellas and wear raincoats on some days, not because one causes the other, but because it rains.
  • Comparing X and Y, X differs from [or is similar to] Y, but neither produces this effect in the other — Temperatures fall and winds rise on the same days, not because one causes the other, but because a cold front passing through causes both.

Is the claim in this article — that signature moves enable greater fluency and which move matters depends on the product in use — causal, correlational, or some mix of both?

Watch for action words that imply cause and effect — words like "lifts," "drives," "leads to." Here's the piece, in its own words:

Each Claude surface rewards a different behavior at the start. We call this the signature move: the gateway behavior that, when present, lifts the other fluency indicators most reliably.

In Chat, the signature move is iterating. Users who refine through follow-up turns show stronger fluency on every other dimension we measure, and users who send one message and leave show almost no critical evaluation at all. Iteration creates the space where other skills develop. Break it down into plain English and pay attention to the verbs...

The signature move [subject] lifts [verb] fluency [object] — a direct causal claim, stated as the thesis.

Then, in the case of Chat specifically, where the signature move is iterating: Users who refine [subject] show [verb] stronger fluency [object] — a correlation stating that those who refine (or iterate) are more fluent than those who leave. No cause is given, only the circumstance under which the pattern was observed — in Chat. Users show some quality (in this case, fluency), but there is no statement alluding to what has induced this quality, only that it tends to occur in Chat. This sentence isn't introducing separate fact that confirms the first sentence even though it may sound that way. It is simply the first sentence restated at a narrower level of specificity — i.e., engaging in iteration in Chat results in stronger fluency than those who don't engage in iteration in Chat. These aren't a claim and its supporting evidence. They're essentially the same claim, once in the abstract and once with detail attached. What may look like supporting evidence is repetition wearing more specific language.

Then iteration [subject] creates [verb] the space where other skills develop [subject] . Creates is another causal verb, but it's used indirectly. It isn't directly producing skill development. It is creating the space for it. While that sounds like an explanation for how iteration leads to stronger fluency, it doesn't actually say how iteration relates to skill development, only that some kind of room gets made for it.

So what is the claim? And is it supported?

Stated plainly, performing a product-specific signature move (e.g., iterating, in Chat) causes a person's overall fluency to increase.

Is it causal, correlational, or both?

The article states its claim as causation (signature moves lead to fluency / X causes Y) but backs it with correlation (users who iterate are more fluent than those who don't / X differs from Y), and never explains the difference. Iterating is itself one of the 11 things counted toward "fluency." So finding that it correlates (X and Y co-occur) with the rest of fluent behaviors is a bit like finding that one ingredient in a soup tastes like soup. And since fluency here just means "good at using Claude," the finding boils down to: people who use Claude well tend to use Claude well — a tautology, not a finding. That's true of almost any skill (the more you use it, the more adept you become) and doesn't need a research report to prove it.

4. If the piece claims something changes "over time," was time actually measured?

This is a subtle one, so it's worth learning to spot. The article says certain skills "grow organically and non-linearly with time and exposure." To show that, you'd need to track the same people over time and watch them individually improve — that's longitudinal data. What the article actually has is a comparison between two different groups at a single point in time: newer users and longer-tenured ones.

That comparison is already broken. There's no way to measure improvement, since no one's change over time was ever tracked — only a snapshot of two different groups. And the two groups start on unequal footing to begin with: new users haven't had time to try most of these behaviors yet, much less become fluent in them. A "new user" snapshot will necessarily look less fluent than an "experienced user" snapshot no matter what. (There's a second, related problem worth a search if you want to dig further: survivorship bias — the possibility that the "long-tenure" group aren't people who got better, but a group that was already better, filtered by who stuck around.)

5. How is the claim tested?

To understand whether the findings are actually meaningful, it is important to check if the method of measurement actually measures what is intended to measure (validity) and whether that measure would produce consistent findings if repeated (reliability).

Take customer satisfaction. A support team might measure customer satisfaction by how quickly a call ends; however, call length can mean different things. A short call could mean the problem got solved fast, or it could mean the agent rushed the customer off the phone without really helping them. In the first case, the customer might be satisfied and in the second the customer may be angry. So call length may not be a valid way to gauge satisfaction.

In this article, fluency — the idea of “getting good at Claude” — is what is being measured. Because it can be defined differently depending on who you ask, the researchers need to operationalize the concept: give it a precise, workable definition, and pick a measure that is concrete, observable, and provides evidence the phenomenon of fluency is actually present. The 11-indicator checklist is Anthropic's measure. It is based on 24 indicators — user behaviors when interacting with AI — that comes from an existing academic framework developed by outside researchers (see https://www.anthropic.com/research/AI-fluency-index and https://aifluencyframework.org/). For this study, Anthropic's research team repurposed the framework as a behavioral classifier for coding conversation transcripts. For each conversation, an AI model checks for the presence of 11 of those 24 indicators – a yes/no judgement call, applied with the same criteria across the entire dataset.

Does this measure measure what it claims to measure? Will the presence of those indicators mean a user is fluent? Technically, no. What it does tell us is which behaviors users commonly engage in. The mere presence of the behavior alone, however, doesn't say whether the behavior leads to an outcome like greater fluency. Did the person actually get what they needed when they engaged in the behavior? Did their code work? Did their report get finished? Were they satisfied with the result? None of that was measured. Second, the article defines fluency as change over time, but this measure only captures a single snapshot. A one-time yes/no checklist can't show whether anyone is developing a skill, no matter how well-built the checklist is.

Could this measure, if repeated under similar conditions, produce the same conclusions? If a behavior like "iteration" is judged by a model reading each conversation, the natural question is: would another rater — a different model, or a human reviewer — looking at the same conversation, and make the same call? If two raters only agree half the time, the measurement is too subjective to trust. The report's "reliability" check shows only that its own model applies its own rules consistently across different days and languages — i.e., the system doesn't produce different results when performing the same task regardless of day or language. It says nothing about whether an independent rater — a completely different system or a human rater — would agree with any given call. That's a different, more basic question, and the report doesn't address it.

Beyond the rater-agreement question, there's a more basic problem: almost nothing about the methods used is made public. The report says a model checked each conversation for the presence of each indicator, but the actual instructions given to that model — the working definition it applied for something like "iteration" — are not included in this piece, and there's no indication of any comparison against human-labeled examples, no accuracy or agreement statistic provided, no account of how ambiguous cases were handled. This is exactly why publishing methods matters: it lets an outside reader check the work, not just take the conclusions on faith. Without the details, there's no way to independently evaluate whether a single "yes" or "no" call means what it claims to.

6. Are concepts clearly defined, and do they hold together — or pull against each other?

This one takes more work, but it's a good habit for anything that proposes a taxonomy — a system of distinct categories. This piece treats "iterating" (its signature move for Chat) and "clarifying the goal" (its signature move for Claude Code and Cowork) as two separate skills. Yet when examined more closely, it's hard to tell them apart: stating a clear goal is often the final step in the process of refining your own thinking (i.e., iterating on the idea) before you sent the chat message, as opposed to after you sent the message with "iterating" — essentially, the same underlying behavior, just timed differently. Before accepting that two labeled things are truly separate, it's worth asking what would be different about them in practice, and whether the difference was tested instead of asumming it.

The same question applies to "Discernment" (judging whether a response is good or complete), which the article says doesn't grow with tenure or practice at all unlike "Description" skills (e.g., revising your goal statement, iterating), which it says develop "organically and non-linearly with time and exposure." But discernment appears to be required for description skills to work: you need to be able to discern whether the response provided was useful or not, so you can revise your goal statement or iterate further if needed. Description and Discernment seem intrinsically linked, to me, so I am confused by the article's claim that Description skills are said to improve with experience, but the judgment embedded in them is not. How can it be true that iterating and clarifying goals improves with tenure but the ability to determine if such tasks are required (discernment) does not?

7. If it's a study of human behavior, does it follow research norms — and if not, why not?

Sorting large amounts of human behavior into labeled categories is an interpretive act, not a neutral measurement. A count, such as "X behavior was present 40% of the time,"" carries no meaning by itself as we discussed in section 5. The article assigns meaning to those counts anyway, treating the presence of a behavior as evidence of "fluency." That inference — higher counts equal fluency — is exactly the kind of claim that needs to be corroborated, not just asserted. Corroborating it means asking questions the count alone can't answer: Why was the behavior present 40% of the time? What conditions in the person's environment might have contributed to that? What was the person thinking, and how might that have motivated them in that situation? Answering questions like these takes interpretive judgment — someone has to decide what's really going on beneath the number. And once you're making a judgment call like that, a further question follows: how do we know the interpretation is trustworthy, rather than just one coder's impression? Qualitative research is the field that has spent decades building tools to answer exactly that question. Inter-rater reliability tests are one such tool. Thick description — providing enough context around a coded instance that a reader can judge the interpretation for themselves — is another. There are others. None of that toolkit, or any substitute for it, shows up in what's published. That doesn't mean the findings are false or inaccurate. It means we, as readers, have no way to know, because nothing about how the judgment calls were made or verified is visible.

Putting it together

Running this piece through those six questions gives you the tools to judge the information for yourself. What those questions reveal is that the article expresses more confidence than its data and methodology can fully support: a central claim that stays somewhat undefined, causal language used in places where the evidence really only shows correlation, a claim about change over time drawn from what looks like a single snapshot, concepts that overlap enough to blur into each other, and methods lacking the detail readers would need to assess the claims.

None of this requires a PhD in research methodology to catch. It requires slowing down at the specific words doing the heavy lifting — "rewards," "grows," "non-linearly," "signature move" — and asking what evidence would actually be needed to back each one up, then checking whether that evidence is present. That habit transfers to just about anything that claims "data shows" or "research finds," not just this one article.