Claude Science on Ozette, part two: reproducing and extending two published findings

In July we put Claude Science on top of Ozette. We connected it to the platform through our MCP server and pointed it at two existing studies: a tuberculosis cohort and a melanoma immunotherapy dataset. That first post was an early demonstration, run on a development copy of the data, meaning the dataset was not processed through our production workflows.

We've since run both studies on Ozette's production platform, with an improved MCP server, and the agent went further than the first time.

In the TB cohort, it reproduced the IFN-γ-independent CD4 T-cell signature of people who resist TB infection, found that only CD4 T cells separate them from latently infected people, and showed that resisters' responses carry fewer functions per cell.

In the melanoma cohort, it reproduced a published biomarker test of anti-PD-1 response exactly, down to the p-value. Then it took the biomarker apart and showed that in this dataset the signal comes from a broader expansion of CD8 effector-memory T cells, not just the specific subset the biomarker describes.

A few findings from the first post change as a result. They're summarized at the end.

Here's what the agent works with. Starting from raw FCS files, with no manual gating, Ozette processes every sample (standardization, preprocessing, quality control and calibration, then Discovery) into high-quality data in which every cell is annotated, marker by marker, and tied to the study design. Through our MCP server, the agent gets the context it needs to reason over that data directly: the studies it's allowed to see, their sample metadata, Ozette Discovery phenotypes, and per-cell annotations it can count in any combination. A scientist can then ask the agent to identify signals in the data that answer questions of interest. Using the MCP connection, the agent can fit statistical models against immunological subsets paired with clinical metadata to answer these questions. In this post, we'll describe two case studies where a scientific driver pairs with an agent supported by the Ozette platform to get insights from their data.

The TB resister cohort

Some household contacts of people with active tuberculosis never test positive for TB infection, even after years of exposure. Their tuberculin skin test (TST) stays negative, and so does their interferon-γ release assay (IGRA), which measures T-cell IFN-γ production in response to TB-specific antigens. Are they uninfected, or responding in a way the tests don't measure?

Lu et al. (Nature Medicine, 2019) studied these "resisters" (RSTR) alongside latently infected controls (LTBI) and found that resisters do respond: they carry antibodies to TB antigens, and CD4 T cells that respond to the IGRA's own antigens by making CD154 (CD40L), IL-2 and TNF without IFN-γ. Ozette holds the intracellular cytokine staining data from that cohort: 42 subjects (22 resisters, 20 LTBI) and 205 samples across five conditions. Those are unstimulated, M. tuberculosis whole-cell lysate, two peptide pools, and SEB as the polyclonal positive control. We asked the agent to investigate the response described in the paper.

The platform records the peptide pools only as "Pool 1" and "Pool 2." The agent identified them from the paper as ESAT6/CFP10, the TB-specific antigens the IGRA uses, and Ag85A/Ag85B/TB10.4, then confirmed which was which from the data: resisters make no IFN-γ in response to Pool 1, as their clinical definition requires.

Checking the assay

Before testing anything, the agent calibrated each functional marker against SEB: a usable marker is low without stimulation and rises under SEB. In CD4 T cells, five of seven meet this criterion: IFN-γ, CD154, CD107a, TNF and IL-2. IL-17a starts high and falls under stimulation, and IL-4 doesn't respond, so both were set aside, leaving 2⁵ = 32 functional subsets per T-cell compartment. Including IL-17a and IL-4 in tests for difference in response changes nothing: every polyfunctional combination that separates the RSTR and LTBI cohorts reduces to a five-marker subset. SEB induces a similar IFN-γ response in both cohorts (p = 0.80), so there's no sign of a global Th1 defect in either group.

The agent also found that IL-17a reads differently by acquisition batch: 48% of unstimulated CD4 cells in one batch against 5.2% in the other. Each batch is balanced by cohort, and every significant cohort result below holds when cohort labels are permuted within batch.

Assay validation. (a) Median functional marker positivity in unstimulated versus SEB-stimulated CD4 T cells: five markers rise, IL-17a falls and IL-4 stays flat. (b) Polyclonal IFN-γ induction is similar in the two cohorts (p = 0.80). (c) Resisters mount no IFN-γ response to ESAT6/CFP10, while LTBI controls respond strongly.
Assay validation. (a) Median functional marker positivity in unstimulated versus SEB-stimulated CD4 T cells: five markers rise, IL-17a falls and IL-4 stays flat. (b) Polyclonal IFN-γ induction is similar in the two cohorts (p = 0.80). (c) Resisters mount no IFN-γ response to ESAT6/CFP10, while LTBI controls respond strongly.

Reproducing the paper's four subsets

The paper names four IFN-γ-negative CD4 subsets that resisters mount against ESAT6/CFP10. Using the Ozette platform, the agent reconstructed each one and tested it against the same subject's unstimulated sample, in resisters alone (21 with paired samples):

CD4 subset (from the paper)Median response, % of CD4p (paired, resisters only)
TNF⁺IL-2⁺CD154⁺IFN-γ⁻0.02811.3 × 10⁻⁴
IL-2⁺CD154⁺IFN-γ⁻0.07341.0 × 10⁻³
CD154⁺ alone0.00525.6 × 10⁻³
CD107a⁺ alone0.06710.026

All four are significant. None of the 16 IFN-γ⁺ combinations responds; the largest median is 0.0004% of CD4 cells. Resisters respond to the IGRA antigens, just not with IFN-γ. That second half is expected, since resisters are selected for a negative IFN-γ test to these same antigens. The concordance of the four IFN-γ-negative responses with the selection criteria provides evidence the agent reproduced real signals in these data.

The paper identified its subsets with COMPASS, a Bayesian model of antigen-specific cytokine combinations. The agent counted the same combinations from Discovery's per-cell annotations and tested them with paired rank tests. Using this alternative analysis route, the agent found the same subsets responded in the same direction as the subsets described in the paper.

The four IFN-γ-negative CD4 subsets the paper attributes to resisters, each tested against the same subject's unstimulated sample. Resisters (blue) respond in all four; LTBI controls (orange) are shown for comparison. Lines join the two samples from one subject.
The four IFN-γ-negative CD4 subsets the paper attributes to resisters, each tested against the same subject's unstimulated sample. Resisters (blue) respond in all four; LTBI controls (orange) are shown for comparison. Lines join the two samples from one subject.

CD4 versus CD8

Both T-cell compartments respond to TB antigens. With lysate, 30 of 32 CD4 subsets and 6 of 32 CD8 subsets rise significantly. But only CD4 separates the cohorts. Across every subset and antigen, 22 comparisons reach cohort-level significance (Benjamini–Hochberg q < 0.05), and all 22 are CD4. LTBI carries the larger response in 20 of them, and TNF appears in 20. The strongest CD4 cohort signal reaches q ≤ 1.5 × 10⁻⁴ against every TB antigen, while no CD8 comparison crosses q = 0.05. SEB separates the cohorts in neither compartment.

CD8 T cells do respond. CD8 IFN-γ⁺ cells rise with lysate in 31 of 39 subjects (p = 8.8 × 10⁻⁵), with no difference between the cohorts (p = 0.75). CD154, which Lu et al. used to mark antigen-specific CD4 cells, behaves exactly as it should: CD154⁺ CD4 cells rise with lysate in 38 of 39 subjects (p = 7.3 × 10⁻¹²) and separate the cohorts (p = 3.4 × 10⁻⁷), while CD154 is essentially absent from CD8 T cells (0.44% even under SEB). At the level of cell types, polyfunctional Th1 cells rise with lysate in all 39 subjects, 10 times more in LTBI than in resisters. Polyfunctional Tc1 cells respond too, with no cohort difference.

Cohort comparison. (a, b) The polyfunctional TNF⁺IL-2⁺CD154⁺ CD4 response to lysate, without and with IFN-γ, is far larger in latent infection. (c) Across the 22 cohort-significant combinations, the LTBI median exceeds the resister median in 20. (d) Strongest cohort signal by compartment: CD4 reaches q ≤ 1.5 × 10⁻⁴ against every antigen, while no CD8 comparison crosses q = 0.05.
Cohort comparison. (a, b) The polyfunctional TNF⁺IL-2⁺CD154⁺ CD4 response to lysate, without and with IFN-γ, is far larger in latent infection. (c) Across the 22 cohort-significant combinations, the LTBI median exceeds the resister median in 20. (d) Strongest cohort signal by compartment: CD4 reaches q ≤ 1.5 × 10⁻⁴ against every antigen, while no CD8 comparison crosses q = 0.05.

Fewer functions per cell

Grouping the responding CD4 cells by how many functions each one carries separates two things: how many cells respond, and how many functions each responding cell has.

Single-function responses don't differ between the cohorts for any antigen. Differences emerge as cells carry more functions: from three functions up, LTBI responses are significantly larger for every antigen, and at four functions they are 12 to 40 times larger than resister responses. For the two peptide pools, the total number of responding cells doesn't differ between cohorts (p = 0.12 for ESAT6/CFP10, p = 0.63 for Ag85/TB10.4). For lysate, the total differs only because of the polyfunctional cells. Among subjects with a measurable response, a median of 20% of a resister's lysate-responding CD4 cells carry two or more functions, against 47% in LTBI (p = 1.1 × 10⁻⁵). The same pattern holds for both peptide pools.

Lu et al. describe resister responses as smaller than those in latent infection, with a lower COMPASS polyfunctionality score that they attribute to the absence of IFN-γ. These data point to something more specific: resisters' single-function responses don't differ from those in latent infection, but they carry far fewer polyfunctional cells. The signal doesn't depend on fine-grained phenotypes, either. Simple two- and three-marker readouts carry it too, and 77 of the 78 tested separate the cohorts. After lysate stimulation, for example, CD154⁺TNF⁺ CD4 cells rise by 0.25% of CD4 in resisters and by 1.57% in LTBI.

Polyfunctionality gradient. (a) Median CD4 response over unstimulated, grouped by the number of functions per cell, for LTBI (solid) and resisters (dashed); asterisks mark cohort q < 0.05. (b) Per-subject fraction of responding CD4 cells with two or more functions.
Polyfunctionality gradient. (a) Median CD4 response over unstimulated, grouped by the number of functions per cell, for LTBI (solid) and resisters (dashed); asterisks mark cohort q < 0.05. (b) Per-subject fraction of responding CD4 cells with two or more functions.

A melanoma immunotherapy biomarker

Greene et al. (Patterns, 2021) found a CD8 effector-memory T-cell phenotype, CD28⁺, HLA-DR⁺ and PD-1⁺, associated with anti-PD-1 response in a Merkel cell carcinoma flow cytometry trial. They then re-analyzed published data from an independent CyTOF dataset in the context of melanoma immunotherapy (Subrahmanyam et al., 2018): they reported that when they tested frequencies of the specific phenotype identified in the Merkel cell carcinoma trial for a difference in response status in the melanoma immunotherapy dataset, the phenotype was more abundant before treatment in patients who went on to respond to pembrolizumab. In full, it is CD8⁺ effector memory (CD45RA⁻CCR7⁻), CD28⁺, CD127⁻CD25⁻ and PD-1⁺HLA-DR⁺.

The published CyTOF dataset has also been re-analyzed in the Ozette platform: 64 patients, each sampled before treatment, unstimulated and after PMA/ionomycin. Each patient received a single drug, either pembrolizumab (anti-PD-1; 21 responders, 19 non-responders) or ipilimumab (anti-CTLA-4; 10 responders, 14 non-responders), with response defined as progression-free survival of at least 180 days.

Reproducing the published test

The agent re-ran the paper's test on the 40 unstimulated pembrolizumab samples: a binomial mixed model with the phenotype counted out of CD3⁺ T cells. It gets an odds ratio of 1.78 and a one-sided p-value of 0.036, which is the value the paper reports. Two-sided, p would be 0.073, so the match also implies the paper's test was one-sided.

Taking the phenotype apart

Since the paper reported testing a specific phenotype in the melanoma dataset, it's natural to wonder which of the markers in the phenotype drive the difference in response status. Since the data on Ozette's discovery platform is highly structured, the agent was able to directly look into this question: it took the phenotype apart one gating step at a time, testing each step against the population above it with the same model. The agent found that the association with response to therapy enters at a single step in the melanoma CyTOF dataset: the split of CD8⁺ T cells into effector memory, which is larger in responders (odds ratio 1.74, one-sided p = 0.025). The CD8 compartment as a whole doesn't differ (p = 0.33). Every condition added after the split, CD28⁺, then CD127⁻CD25⁻, then PD-1⁺HLA-DR⁺, has an odds ratio near 1 against the population above it (p = 0.48, 0.59 and 0.57). In other words, responders have more effector-memory cells, but the PD-1⁺HLA-DR⁺ share of its immediate parent is the same in both groups.

If PD-1⁺HLA-DR⁺ cells were the response-specific part of the phenotype in the melanoma CyTOF dataset, they should stand out from similar phenotypes that are either PD-1 negative, HLA-DR negative, or double negative for those markers. They don't. All four PD-1 × HLA-DR quadrants rise to a similar degree as a fraction of CD3⁺ T cells, and none differs as a fraction of its parent under the gating strategy derived by the agent.

The p-value is also driven by a subset of subjects. A rank-based test gives one-sided p = 0.092, and removing any one of 10 of the 40 patients pushes the result above 0.05. Age doesn't explain any of this - the agent checked for an age effect.

In this melanoma dataset, then, the "PD-1⁺HLA-DR⁺CD28⁺" label describes cells that come along with a broader expansion of effector-memory CD8⁺ T cells in responders, rather than a response-specific subset. One question this analysis can't answer: the Merkel cell carcinoma discovery dataset distinguished PD-1 dim from PD-1 bright cells, while the gating in the melanoma CyTOF dataset has a single PD-1 cut-off and so only identifies cells that are either PD-1 positive or PD-1 negative. Therefore one question that is raised from this agentic re-analysis is: do PD-1 bright cells identify response-specific subsets within the CD8 effector memory compartment in the melanoma dataset?

Melanoma, unstimulated pembrolizumab samples. (a) The paper's test reproduces: odds ratio 1.78, one-sided p = 0.036, the published value. (b) The association enters at the naive-to-effector-memory split; every later gating step has an odds ratio near 1 against its parent. (c) Raw data: effector memory expands in responders, while the PD-1⁺HLA-DR⁺ share of its parent is flat. (d) All four PD-1 × HLA-DR quadrants rise alike as a share of CD3⁺ T cells, and none differs as a share of its parent. (e) Removing any one of 10 of the 40 patients pushes the one-sided p-value above 0.05; a rank test gives 0.092.
Melanoma, unstimulated pembrolizumab samples. (a) The paper's test reproduces: odds ratio 1.78, one-sided p = 0.036, the published value. (b) The association enters at the naive-to-effector-memory split; every later gating step has an odds ratio near 1 against its parent. (c) Raw data: effector memory expands in responders, while the PD-1⁺HLA-DR⁺ share of its parent is flat. (d) All four PD-1 × HLA-DR quadrants rise alike as a share of CD3⁺ T cells, and none differs as a share of its parent. (e) Removing any one of 10 of the 40 patients pushes the one-sided p-value above 0.05; a rank test gives 0.092.

How the agent got there

None of this required a bioinformatician writing custom code for each step. Because every cell in Ozette is annotated, marker by marker, the agent could:

Corrections to the first post

The first post ran on a development copy of these studies (meaning the data were not processed through our production pipelines; it was an old run), and its marker calls don't pass the calibration check described above: five of the seven functional markers read positive on 71–100% of unstimulated CD8 cells, and the median sample had 20,521 cells called both CD4⁺ and CD8⁺. Three TB findings change as a result:

One conclusion holds, on firmer ground: CD154 marks the antigen-specific response, with CD154⁺ CD4 cells rising in 38 of 39 subjects.

In the melanoma case study, the first post described the patients as receiving combination checkpoint blockade and analyzed both treatment arms together. Each patient received a single drug, and the published test belongs to the pembrolizumab arm, where it reproduces at the paper's one-sided p = 0.036, as described above. The UMAP figure in the first post came from the development copy, so treat it as an illustration of the workflow rather than of the biology.

Limitations

Both studies are re-analyses of published cohorts. The TB study has 42 subjects. Its screens are corrected for multiple comparisons within each compartment and antigen, the targeted tests of the paper's four subsets are reported unadjusted, and the polyfunctionality finding is a hypothesis worth testing in independent data. The melanoma comparison has 40 patients and uses the dataset the paper itself relied on, so it reproduces the paper's test rather than validating the biomarker independently.

Final thoughts

Flow cytometry archives hold years of immune data that was analyzed once. With every cell annotated and tied to the study design, an agent can reproduce a published analysis, extend it, and test what a published biomarker actually measures, in a single session rather than a new project. Is the resister response smaller, or just less polyfunctional? Is a biomarker from one trial and one technology a specific subset, or a marker of something broader? With the data in this form, those are some of the questions you can ask. To be clear - you are not limited to reanalysis either. An agent can analyze data de novo with the right context and questions driving its work.

This is available today. If you already work with Ozette, your processed studies are ready to connect to Claude Science, or another MCP-compatible agent your organization uses, through our MCP server. If you don't, we can load and process you historical flow archive onto the Ozette platform. Get in touch at bd@ozette.com, and join us on Thursday, October 15, at 9:00 AM PT for our webinar, Agentic Science Meets Structured Immune Data, where we'll walk through both studies. Register here.


Figures in this post were generated by Claude Science operating over the Ozette MCP server. Statistical values are as computed in-session. References: Lu LL et al. IFN-γ-independent immune markers of Mycobacterium tuberculosis exposure. Nat Med 2019;25:977–987. Subrahmanyam PB et al. Distinct predictive biomarker candidates for response to anti-CTLA-4 and anti-PD-1 immunotherapy in melanoma patients. J Immunother Cancer 2018;6:18. Greene E et al. New interpretable machine-learning method for single-cell data reveals correlates of clinical response to cancer immunotherapy. Patterns 2021;2:100372.

← All posts