Roseann white discussing clinical trial statistics and FDA Review

In clinical trials, statistics influence almost every major decision.

From endpoint selection and sample size to regulatory strategy, product approval, and post-market evidence, biostatistics shapes how clinical trials are designed, analyzed, and interpreted.

But for many clinical research professionals, statistics can feel intimidating. It can feel too academic, too technical, or disconnected from the practical decisions that sponsors, CROs, investigators, and clinical operations teams need to make.

Roseann White is a retired biostatistician with more than 30 years of experience across research, development, manufacturing, and clinical research. Her work has included medical device and diagnostic trial designs, statistical analysis plans, FDA interactions, regulatory submissions, and FDA advisory panel support.

In this episode of the Clinical Trial Podcast, Roseann makes biostatistics practical.

She explains what statistical significance does and does not mean, why clinically meaningful results matter, how to think about superiority and non-inferiority, and why sensitivity and specificity are connected rather than separate concepts.

Roseann also discusses why statisticians should be involved early in trial strategy, how financial constraints influence study design, what happens when FDA reviewers disagree with a methodology, and why non-statisticians should ask, “What can go wrong?” before the trial is already underway.

This conversation is for anyone involved in designing, conducting, analyzing, reviewing, or interpreting clinical trials.

What You’ll Learn

In this episode, you will learn:

  • What “statistically significant” really means, and why it does not guarantee that a product works
  • Why clinical meaning matters just as much as statistical significance
  • How to think about superiority and non-inferiority in practical terms
  • Why non-inferiority can be useful, and when it may be the wrong design choice
  • How sensitivity and specificity are connected in diagnostic testing
  • Why ROC curves help show the relationship between sensitivity and specificity
  • Why statisticians should be involved early in clinical trial strategy
  • How a statistical analysis plan can help avoid confusion later in a trial or registry
  • Why endpoint selection should begin with the clinical question, not just the statistical method
  • How financial constraints can influence sample size, power, and study design decisions
  • Why FDA disagreements should be handled with clarity, humility, and strong explanation
  • Why statistical methods should be understandable to both FDA statisticians and clinical reviewers
  • What non-statisticians should ask earlier in the trial design process
  • Which resources Roseann recommends for people who want to understand statistics better

About the Guest

Roseann White is a retired biostatistician with more than 30 years of experience providing statistical guidance across research, development, manufacturing, and clinical research.

Her work has included medical device and diagnostic trial design, statistical analysis plans, FDA interactions, regulatory submissions, and FDA advisory panel support. She has also worked with companies on clinical trial designs and analyses for regulatory approval and reimbursement.

In retirement, Roseann is focused on helping professionals build a more intuitive understanding of statistics through writing, social media, and volunteer efforts. In the interview, she described this work under the banner of “Humble Statistician.”

Selected Links from the Episode:

Books:

Show Notes

[0:00] Clinical Trial Podcast Opening

  • The podcast’s mission is to help clinical research professionals develop into effective leaders.
  • Host Kunal Sampat introduces the Clinical Trial Podcast.

[0:20] Kunal’s Introduction and Roseann White’s Background

  • Statistics influences endpoint selection, sample size, data interpretation, budgets, regulatory strategy, approval, and patient access.
  • Biostatistics can feel intimidating or disconnected from the day-to-day work of conducting clinical trials.
  • Roseann White has more than 30 years of experience across research, development, manufacturing, and clinical research.
  • Her experience includes medical device and diagnostic trial design, statistical analysis plans, FDA interactions, regulatory submissions, and advisory panel support.
  • The episode will make statistical concepts more practical for clinicians, sponsors, operations teams, and other non-statisticians.

[3:48] Roseann’s Goal to Make Statistics More Accessible

  • Roseann plans to write blogs that explain complex statistical issues in a straightforward manner.
  • Her goal is to give people an intuitive understanding of the methods affecting their work.
  • Professionals should not feel that their only option is to trust a statistician without understanding the reasoning.
  • Better statistical literacy enables study-team members to participate more effectively in trial-design decisions.

[4:44] What Statistical Significance Does and Does Not Mean

  • A statistically significant result does not automatically prove that a product works.
  • The significance threshold is a decision boundary used to control the risk of concluding that a treatment effect exists when it does not.
  • A result just below 0.05 and one just above 0.05 should not be interpreted as completely different scientific realities.
  • Statistical significance reflects uncertainty and risk; it is not a guarantee that a result will repeat in another study.
  • Different programs may use more stringent or more permissive thresholds depending on the seriousness of the decision and the acceptable level of risk.

[7:33] Why Five Percent Became the Conventional Threshold

  • Roseann recounts the historical story of statisticians Ronald Fisher and Jerzy Neyman.
  • Fisher published compact statistical tables using thresholds such as 0.05, 0.01, and 0.001.
  • Those tables were easier for researchers to use than larger and more complicated alternatives.
  • The 5% threshold became a practical convention rather than a universal law of science.
  • Modern software has replaced the need to consult printed statistical tables, but the convention remains widely used.

[9:19] Superiority and Non-Inferiority in Practical Terms

  • Roseann uses Nike sneakers and less expensive alternatives to explain non-inferiority.
  • A non-inferiority study asks how close an alternative must be to an established treatment to remain acceptable.
  • Non-inferiority does not mean that two products are exactly equivalent.
  • Superiority asks whether one product is meaningfully better than another.
  • A superiority claim should represent a difference that matters clinically, not merely a tiny difference that becomes statistically significant because the sample is extremely large.
  • Roseann emphasizes that almost any small difference can become statistically significant with a sufficiently large sample.

[13:52] Choosing Between Superiority and Non-Inferiority

  • Non-inferiority may be appropriate when a new product is expected to perform similarly to an existing treatment but offers another benefit.
  • Additional benefits might include lower cost, greater accessibility, improved safety, convenience, or a less invasive procedure.
  • Roseann discusses the choice between open surgery and a percutaneous intervention as an example of a treatment that may be attractive even if the efficacy results are not clearly superior.
  • Smaller companies may select non-inferiority because they cannot afford the sample size needed to establish superiority.
  • A study may first test non-inferiority and then proceed to superiority when that testing sequence is prespecified.
  • A product still needs a compelling clinical or practical advantage to justify a non-inferiority strategy.

[17:30] How Statistical Margins Affect Sample Size

  • Superiority does not always require a larger sample than non-inferiority.
  • Sample size depends on the expected treatment effect, variability, event rates, and the selected superiority or non-inferiority margin.
  • A narrow non-inferiority margin can require more participants than a study designed to detect a large superiority effect.
  • Teams should avoid making general assumptions about sample size based only on the name of the study design.

[18:47] Clinical Meaning Must Drive the Statistical Design

  • Statisticians help translate a clinically meaningful difference into hypotheses, margins, and sample-size requirements.
  • The statistician should not independently decide what magnitude of benefit matters to patients or clinicians.
  • Clinical experts must determine how different two products need to be before the difference would influence treatment selection.
  • The same clinical reasoning is needed when deciding how much worse a new product could be while remaining an acceptable alternative.
  • Roseann summarizes the relationship by stating that science should drive statistics—not the other way around.
  • Statistical significance does not automatically make a difference clinically meaningful.

[20:33] When Non-Inferiority Is the Wrong Choice

  • Non-inferiority may be the wrong primary design when the product is genuinely expected to be superior and a superiority study is feasible.
  • A trial powered only for non-inferiority may lack enough participants to establish a superiority claim.
  • Even when the observed results favor the new treatment, the study may miss a potentially valuable labeling claim.
  • Roseann discusses a coronary stent program in which additional post-approval work was needed to investigate superiority in selected subgroups.
  • Early strategic discussion can prevent sponsors from having to conduct another study to obtain a claim that could have been addressed initially.

[22:55] Sensitivity and Specificity Are Connected

  • Sensitivity and specificity should not be treated as independent measures.
  • Changing a diagnostic threshold generally changes both measures.
  • Increasing the ability to identify one type of result may increase the number of errors in the other direction.
  • The right balance depends on the clinical consequences of false-positive and false-negative findings.
  • The study team must understand which type of error presents the greater risk for the intended use of the diagnostic.

[25:28] When the Statistician Should Join Trial Strategy

  • The statistician should be involved when the overall product-approval strategy is being developed.
  • Waiting until the pivotal or Phase III study is too late to obtain the full benefit of statistical guidance.
  • Earlier studies should build evidence that supports assumptions for later-stage trial design.
  • Early involvement can prevent teams from collecting the wrong data, overlooking essential endpoints, or creating avoidable analysis problems.
  • A statistician can help determine whether preliminary data are adequate to support the proposed pivotal study.

[27:56] Statistical Analysis Plans for Registries

  • Registries also need clear objectives and a statistical analysis plan.
  • Teams should decide whether a registry is descriptive, hypothesis-generating, or intended to support a formal hypothesis test.
  • A registry designed to answer a specific question may require a planned sample size or number of events.
  • The timing of analyses should be based on when enough information will be available—not simply on an arbitrary calendar date.
  • Without a clear plan, teams may repeatedly examine the data and later disagree about what the findings mean.

[33:45] Multiple Testing, P-Values, and Hypothesis Generation

  • Testing many variables increases the chance of observing apparently significant findings by chance.
  • A p-value is most interpretable when it comes from a properly designed and powered hypothesis test.
  • Exploratory registry analyses may be valuable, but their findings should generally be treated as hypothesis-generating.
  • A clinically meaningful pattern may deserve further study even when its p-value does not cross the conventional threshold.
  • Conversely, a p-value below 0.05 should not become the only reason to pursue a finding.
  • Roseann references the American Statistical Association’s statement cautioning against using p-values as the sole measure of scientific importance.
  • Scientific rationale should determine which questions deserve further investigation.

[34:47] Declaring the Registry’s Intentions Up Front

  • Teams should specify the registry’s primary objectives before analyzing the data.
  • Additional analyses can be identified as exploratory and used to generate future hypotheses.
  • Emerging findings may justify changing enrollment to obtain more information about a particular population.
  • Prespecification gives researchers and regulators a shared framework for interpreting what was planned and what was discovered later.
  • Clear documentation reduces arguments about whether the analysis was conducted as intended.

[35:09] Revising Plans for Multi-Year Registries

  • A statistical analysis plan may be revised when new scientific questions emerge during a long-term registry.
  • Revisions should be documented clearly, including when the change occurred and why it was made.
  • The team should distinguish questions developed before examining relevant outcomes from analyses prompted by observed data.
  • Long-term follow-up may require different endpoints from those used for the original approval objective.

[35:55] Extension Statistical Analysis Plans

  • Roseann discusses an extension statistical analysis plan used for longer-term follow-up.
  • The original plan may focus on the endpoint needed for approval at one year.
  • A separate extension SAP can address outcomes that become more important over three, four, or five years.
  • The long-term primary endpoint does not necessarily need to be identical to the initial endpoint.
  • The extension plan must still address multiplicity and control the risk of false-positive conclusions.
  • PARTNER 3 is discussed as an example of a program with extended patient follow-up.

[37:20] ROC Curves and Diagnostic Tradeoffs

  • ROC stands for receiver operating characteristic.
  • An ROC curve displays the relationship between sensitivity and specificity across possible diagnostic thresholds.
  • It helps the study team identify a point at which both measures are acceptable for the product’s intended use.
  • A perfect diagnostic would not require a meaningful tradeoff, but real diagnostic products generally do.
  • Roseann uses colorectal cancer screening to illustrate how the clinical cost of missed cases and false alarms influences threshold selection.

[39:34] Where to Start When Building a Statistical Analysis Plan

  • The first question is: What is the main objective of the study?
  • The statistician then asks how the product is expected to be better and what outcome would demonstrate that benefit.
  • The clinical team and statistician define the primary endpoint and the magnitude of difference that would be clinically meaningful.
  • The statistician can show how different assumptions and effect sizes affect the required sample size.
  • This discussion produces the primary objective, primary hypothesis, and primary analysis.

[41:00] Competing Risks and Events Unrelated to the Product

  • The primary endpoint must be examined for events that could interfere with its interpretation.
  • For example, death may prevent a patient from experiencing a later hospitalization endpoint.
  • The analysis must account for outcomes that alter whether the primary event can occur or be observed.
  • External events such as the COVID-19 pandemic can change healthcare utilization and invalidate assumptions made before the trial.
  • A robust analysis plan anticipates factors that could affect the endpoint even when they are unrelated to the investigational product.

[45:27] Missing Data, Subgroups, and Sensitivity Analyses

  • The SAP should explain how missing data and patient withdrawals will be handled.
  • Planned subgroup analyses should be identified in advance.
  • Sensitivity analyses help determine whether the primary conclusion remains stable under different reasonable assumptions.
  • Secondary and exploratory endpoints can be included after the primary endpoint has been fully defined.
  • Roseann describes the primary endpoint and its possible complications as the central focus of SAP development.
  • The objective is to obtain the most robust and defensible answer the trial can provide.

[46:10] Knowing When Specialized Statistical Expertise Is Needed

  • A study team may need an external specialist when the proposed method is outside its statistician’s area of expertise.
  • Roseann suggests evaluating whether the statistician can explain the method in terms the clinical team understands.
  • A person who understands the details and limitations of a method should be able to communicate it clearly.
  • The ability to simplify an explanation without misrepresenting the method reflects deep knowledge.
  • Roseann is comfortable discussing Bayesian concepts but would refer a fully Bayesian trial design to a Bayesian specialist.
  • Recognizing the limits of one’s expertise is part of responsible statistical leadership.

[48:22] How Financial Constraints Shape Statistical Choices

  • One of the first practical questions is how large a study the sponsor can afford.
  • The available budget defines the design space within which the statistician must work.
  • Smaller companies often need creative but statistically defensible ways to reduce sample size.
  • Some proposed studies cannot be made adequately informative within the available budget.
  • Sponsors may reject a more sophisticated design because it seems complicated and instead propose a simpler approach to FDA.
  • Roseann notes that FDA can function as a statistical guardrail by rejecting plans that are not sufficiently strong.

[51:00] The Risks of Underpowering a Trial

  • Statistical power is the probability of detecting a prespecified treatment effect when that effect is present.
  • A study designed with 80% power still carries a meaningful probability of failing to detect the assumed effect.
  • Some sponsors accept lower power or use optimistic assumptions to reduce the apparent sample-size requirement.
  • Those choices can produce a trial that fails even when the product has some benefit.
  • An underpowered study may also lack enough information for regulators to interpret secondary endpoints or determine why the primary analysis failed.
  • For a startup dependent on approval, the need to conduct another trial can create a major financial and strategic setback.

[54:10] Adaptive Designs as an Early-Warning and Financing Tool

  • Adaptive sample-size re-estimation allows the team to reassess whether the original sample size remains adequate.
  • Enrichment designs may focus later enrollment on a subgroup that appears more likely to benefit.
  • Other adaptations can involve randomization or, in appropriate circumstances, an endpoint.
  • Adaptations must be prespecified and implemented using appropriate methods to protect trial validity.
  • An interim re-estimation may show a small company that only a limited number of additional participants are needed.
  • The resulting information can help the company approach investors with a clearer estimate of the funding needed to finish the study.
  • Adaptive design can therefore support both statistical risk management and business planning.

[56:01] Managing Mid-Trial Sample-Size Changes

  • Kunal describes a trial whose planned enrollment was reduced from approximately 2,600 participants to approximately 1,900.
  • Business constraints can create pressure to modify a study after enrollment has begun.
  • A change is easier to manage when the potential adaptation was anticipated and prespecified.
  • Without a planned mechanism, the change may require extensive discussions among statistical, clinical, regulatory, and business stakeholders.
  • Modern adaptive methods provide more structured ways to address these situations.

[57:25] Choosing an Endpoint When Everything Matters

  • A study may have several clinically important outcomes, including mortality, myocardial infarction, and rehospitalization.
  • The team can establish a natural hierarchy that gives the most consequential outcomes priority.
  • Roseann discusses the win-ratio method for comparing patients across a hierarchy of outcomes.
  • Patients in the treatment and control groups are compared first on the highest-priority outcome.
  • When that outcome does not determine a winner, the comparison proceeds to the next outcome.
  • This method can preserve clinical priorities that may be obscured by a conventional composite endpoint.
  • Roseann also references earlier work by Finkelstein and Schoenfeld.

[59:36] How to Respond When FDA Disagrees with the Methodology

  • Telling an FDA reviewer that the reviewer is wrong is rarely an effective strategy.
  • The first step is to understand the reviewer’s concern and determine whether the alternative method would materially weaken the study.
  • When FDA’s preferred method produces an acceptable analysis, using it may be the practical choice.
  • The team should distinguish disagreement about the clinical endpoint from disagreement about the statistical technique.
  • A complex or Bayesian method must be explained clearly enough for FDA to understand exactly what is being proposed.
  • The SAP should be written for both statistical reviewers and clinical reviewers.
  • A statistically sophisticated document that the clinical reviewer cannot understand may still create regulatory difficulty.

[1:03:20] Knowing the FDA Reviewer and Writing for More Than One Audience

  • The reviewer’s clinical background and prior experience can influence the questions raised during review.
  • Roseann describes a percutaneous product reviewed by a surgeon who strongly favored surgery.
  • The pivotal trial did not initially secure approval through the intended comparison.
  • Approval was ultimately supported in a population for whom surgery was not a viable option.
  • Understanding the reviewer’s concerns earlier could have helped the sponsor design a more direct and efficient development strategy.
  • Statistical and regulatory communication must address the assumptions and concerns of the actual audience.

[1:05:44] Reviewer Turnover and Nonbinding FDA Feedback

  • FDA review teams can change between a pre-submission meeting and a later marketing submission.
  • Pre-submission feedback may be nonbinding, which creates uncertainty when a new team takes a different position.
  • Sponsors should document how previous decisions were reached and explain the consequences of changing the agreed approach.
  • The team should determine whether a compromise—such as a revised margin or analysis—can address the new concern without making the study infeasible.
  • When the requested change cannot be executed reasonably, escalation may be necessary.
  • Roseann recounts a situation in which FDA changed its view of the primary endpoint after the study had already been completed.
  • Clear records, transparent communication, and a willingness to negotiate are essential in these situations.

[1:08:59] Questions Non-Statisticians Should Ask Earlier

  • The first question non-statisticians should ask is: “What can go wrong?”
  • Study teams should ask how the SAP will handle deviations from the assumptions underlying the design.
  • ICH statistical guidance emphasizes planning for intercurrent events and other events that affect interpretation of treatment outcomes.
  • The second question is: “What happens if the trial fails?”
  • Teams should understand what information might still be available and whether any prespecified adjustment or follow-up strategy exists.
  • Clinical development professionals are naturally optimistic about their products, but responsible planning also requires considering failure scenarios.
  • Adaptive studies can provide prespecified options for responding when assumptions are not met.

[1:11:26] Trusted Resources for Understanding Statistics

  • Roseann recommends Clinical Trials: A Practical Approach by Stuart J. Pocock.
  • She praises Pocock’s ability to explain technical concepts clearly without making the audience feel less capable.
  • His articles and lectures can help readers understand how to evaluate papers and interpret statistical methods.
  • For people with little statistical background, she recommends The Cartoon Guide to Statistics by Larry Gonick and Woollcott Smith.
  • The illustrations provide an accessible introduction to fundamental concepts before readers move to more advanced material.

[1:13:31] How Statisticians Stay Current

  • Statisticians should regularly speak with other statisticians working in the same specialty.
  • Specialty-specific collaboration is particularly important because methods and regulatory expectations differ across product areas.
  • Roseann highlights the annual FDA/AdvaMed Medical Device Statistical Issues Conference.
  • The meeting gives industry statisticians opportunities to learn about practical methods and work directly with FDA colleagues.
  • Professional relationships can improve both technical learning and regulatory communication.

[1:15:00] Taking on Difficult Problems

  • Roseann keeps her skills sharp by working on problems that require her to learn new methods.
  • Difficult questions motivate statisticians to search the literature and test different analytical approaches.
  • She describes investigating a problem after a principal investigator said that the proposed work could not be done.
  • Her research identified viable approaches, after which specialists were engaged to complete the work.
  • The resulting study used a manageable sample size and ultimately contributed to an approved product.
  • Challenging assignments can be one of the most effective forms of continuing education.

[1:16:08] Books, Lectures, and Visual Learning

  • Roseann again recommends Clinical Trials: A Practical Approach and The Cartoon Guide to Statistics.
  • She believes statistics is often best taught through a combination of visual presentation and verbal explanation.
  • A written reference to a figure may not communicate the concept as effectively as a skilled lecturer walking through the visual.
  • Recorded lectures and educational videos can therefore be especially useful.
  • Stuart Pocock’s lectures are among the resources she recommends.

[1:18:02] Frequentist and Bayesian Interpretations

  • Roseann refers to a published paper that offers a different way of interpreting frequentist results.
  • The paper’s title is not identified during the interview.
  • People often interpret a frequentist confidence interval as the probability that the parameter lies within that interval.
  • Roseann explains that the conventional frequentist interpretation is based on repeated sampling.
  • Bayesian methods can make direct probability statements about a parameter under the specified model and prior assumptions.
  • She mentions predictive values as a way to communicate information that may feel more intuitive to decision-makers.
  • The discussion illustrates why statisticians must explain exactly what a statistical statement does—and does not—mean.

[1:20:19] When a Statistically Correct Decision Still Leads to a Bad Outcome

  • A decision can follow the prespecified statistical rules and still lead to an undesirable clinical or business result.
  • Roseann describes a program in which a potential safety signal was not statistically significant at approval.
  • Additional real-world use eventually produced enough evidence to make the safety concern clearer.
  • The product’s use had to change after the post-market evidence accumulated.
  • Rare safety events may not be detectable in the sample size used for a pivotal trial.
  • Post-approval studies and post-market surveillance remain important even when the original study was analyzed correctly.
  • Statistical decision rules do not remove the need for continuing clinical judgment and evidence collection.

[1:22:07] Humble Statistician and Roseann’s Closing Thoughts

  • Roseann is retired and is not seeking to build another consulting practice.
  • She wants to use her experience to help others understand statistics more intuitively.
  • Her educational work is being developed under the name “Humble Statistician.”
  • She hopes to share practical explanations through blogs, social media, and volunteer activities.
  • Kunal thanks Roseann for making complex statistical topics approachable.

[1:23:45] Episode Closing

  • Kunal closes the conversation and thanks listeners for joining the Clinical Trial Podcast.

Major Themes

  • Statistical significance must be interpreted in clinical and scientific context. A p-value does not prove that a product works, and a statistically significant result may not represent a clinically meaningful benefit. Clinical objectives and patient relevance must drive the statistical design.
  • Early collaboration and prespecification make studies more defensible. Clinicians and statisticians should work together before the pivotal trial to define objectives, endpoints, assumptions, margins, missing-data strategies, and contingency plans. Prespecified flexibility—including adaptive designs and extension SAPs—can help teams respond to emerging information while preserving trial validity.
  • Budget and regulatory realities are part of statistical strategy. Financial constraints, FDA feedback, reviewer backgrounds, and review-team turnover can materially affect study design. Clear explanations, realistic power assumptions, careful documentation, and planning for failure protect both the development program and the patients who may depend on it.

 

Selected Quotes

  • “I can make anything statistically significant with a large enough sample size.”
  • “Statistics doesn’t drive science; science drives the statistics.”
  • “Only somebody who fundamentally understands the details and the ins and outs of a method can explain it simply to somebody else.”

Audience Question

Which statistical decision has created the most uncertainty on a clinical trial you worked on—endpoint selection, sample size and power, missing data, or planning for what could go wrong? 

Leave a Reply

Your email address will not be published. Required fields are marked *

Share On Facebook
Share On Twitter
Share On Linkedin
Share On Pinterest
Contact us