OpenSurgery OpenSurgery OpenSurgery OpenSurgery

Surgical Audit and Research

Summary

  • Audit and research answer different questions: audit measures practice against a standard that already exists, research generates the standard. Audit is a formal process requiring a structure, and its defining feature is that it closes, the seventh step of the cycle is to reaudit [1].
  • This page covers the audit cycle, the multicentre snapshot audit, how a research question becomes a project, evidence-based surgery and the Cochrane Collaboration, and the IMRAD structure for writing it up.
  • The motivating observation is blunt: sufficient evidence to justify routine thrombolysis for myocardial infarction existed years before the randomised trials that made it acceptable, no one had gathered the available information together [1].

Definition

Research is designed to generate new knowledge, and might involve testing a new treatment or regimen [1].

Evidence-based surgery is a move to find the best ways of managing patients using clinical evidence from collected studies [1].

A multicentre snapshot audit is one in which many collaborators across multiple hospitals prospectively collate anonymised patient-level data for a specific condition, presentation or intervention over a short period, normally around 6 to 8 weeks [1].

Evidence-based medicine as Schwartz defines it

Evidence-based medicine (EBM) is "the conscientious, explicit and judicious use of current best evidence in making decisions about treating individual patients." The term was coined by Gordon Guyatt in 1991, and the framework grew out of work at McMaster University in the latter part of the 20th century; its roots lie in the emergence of clinical epidemiology as a field in 1938, which shifted attention from descriptions of individual patients to trends affecting populations [2].

EBM rests on three epistemological principles [2]:

1. Not all evidence is created equal, practice should be based on the best available evidence. 2. The pursuit of truth is best accomplished by evaluating the totality of the evidence, not by selecting evidence that favours a particular claim. 3. Clinical decision-making requires consideration of patients' values and preferences.

A clinical guideline, in the Institute of Medicine's definition, is a set of "statements that include recommendations, intended to optimize patient care, that are informed by a systematic review of evidence and an assessment of the benefits and harms of alternative care options" [2]. Guidelines often represent the highest level of applied clinical evidence, but they vary in quality just as individual studies do [2].

Pathophysiology

The problem evidence-based surgery exists to solve is variation. Surgical practice has been considered an art: ask 50 surgeons how to manage a patient and you will probably get 50 different answers, and there is so much clinical information available that no surgeon can know it all [1].

Why surgery lags: the scarcity of trials and the roots of bias

  • Surgery adopted EBM later than the nonsurgical specialties, chiefly because the "gold standard" randomised controlled trial (RCT) is hard to run on operations.
  • A MEDLINE analysis from 1966 to 2000 found that only 15.1% of 134,689 RCTs evaluated a surgical topic; in the 1990s surgical RCTs made up only 7% of articles in surgical journals, most of which were retrospective studies and case series; and by 2003 the relative frequency had fallen to 3.4% of all publications [2].
  • Most surgical practice therefore still rests on retrospective reviews, nonrandomised trials and expert opinion, and the barriers, standardising clinical presentation, accounting for variation in operative technique, and the difficulty of blinding, remain substantial [2].

The dangers of cognitive bias were catalogued by Francis Bacon in 1620 as four "idols": idols of the tribe (errors common to human nature), idols of the cave (individual personality and preferences), idols of the marketplace (confusion arising from language, and scientific words taking a meaning different from common usage) and idols of the theatre (following academic dogma without questioning) [2]. Their modern descendants in the biomedical literature are [2]:

BiasMechanism
Publication biasPublishers are incentivised to accept positive results
Prevailing-field biasSupport for entrenched opinions
Citation biasTendency to cite positive studies
Time-lag biasDelay in reporting negative results
Reporting biasEmphasising positive over negative results

Schwartz's point is that these are the individual sources of cognitive bias "writ large across an entire community", and that EBM is an attempt to codify the process of interpreting experience so that a field can advance scientifically despite them [2].

Clinical features

Audit comes in two forms in common practice (single-site local audits, and multisite regional, national or international audits) and both are designed to improve quality of care [1].

In an ideal world the two feed each other: local audits identify needs closest to the patient, which are then investigated in larger multisite audits [1]. Hospital topics are often identified at departmental morbidity and mortality meetings, and the reporting process may identify a possible national issue, prompting a national or international audit delivered by local surgical teams [1].

Etiology

The change an audit demands may sit at any level. Sometimes it is at the level of the individual, sometimes the team; sometimes the only appropriate action is change at institutional level such as a new antibiotic policy, at regional level such as provision of a tertiary referral centre, or at national level such as a screening programme or health education campaign [1].

Failure to involve others is one of the commonest reasons a good research project fails. Only the smallest single-centre project can be delivered by an individual working alone; almost any project worth doing needs a team, both for its skills and, more importantly, for the momentum required to keep pushing the project through when the inevitable hurdles are met [1].

Diagnosis

The literature is searched through defined databases [1]:

DatabaseCoverageAccess
PubMedOver 25 million citations for biomedical literature from MEDLINE, life science journals and online booksFree
PubMed CentralFull-text archive of biomedical and life sciences literature at the US NIH National Library of MedicineFree digital archive
EMBASEExtensive coverage of peer-reviewed biomedical literature with indexing and search toolsSubscription
CINAHLCumulated index to nursing and allied health literatureSubscription
Cochrane LibraryCochrane Reviews, prepared and updated by a global independent network, free from commercial sponsorshipFree

The Cochrane Library includes a database of systematic reviews, reviews of surgical effectiveness, and a register of controlled trials [1].

Framing the search: PICO

MEDLINE via PubMed now holds over 26 million citations, and unstructured searching of it can be overwhelming. Schwartz recommends framing the clinical question in the PICO format [2]:

  • Patient or population, the specific group about whom the question is asked.
  • Intervention, the treatment or technique of interest (a procedure such as "laparoscopic appendectomy", or an exposure such as "smoking").
  • Comparison, the alternative treatment ("open appendectomy", "observation").
  • Outcome, "mortality", "operative time", "wound infection".

There is a trade-off between specificity and breadth, and for clinical decision-making it is generally better to be as precise as possible, chaining terms with "AND", for example (distal pancreatectomy) AND splenectomy AND (splenic preservation) AND morbidity [2].

Classifying what the search returns

Because RCTs are rare in surgery, the surgeon must be familiar with the alternative study types and their weaknesses [2]:

  • Meta-analysis, combines similarly published data to increase statistical power beyond any single study; interstudy heterogeneity (methods, population, endpoints) must be limited, and inclusion of inappropriate studies or mislabelling of a meta-analysis can produce inaccurate conclusions. Look here first when no clinical guideline exists.
  • Systematic review, uses the same standardised search and appraisal methods to reduce bias, but does not pool the results quantitatively, so it is usually regarded as weaker evidence than a meta-analysis. Systematic reviews are nonetheless the most cited type of study and essential to guideline development; applied in a timely fashion they have changed practice, for example the move to early postoperative enteral feeding over parenteral nutrition to prevent sepsis.
  • Cross-sectional study, exposures and outcomes measured at a single point in time and prevalence compared between exposed and unexposed; many exposures and outcomes can be measured at once, but no temporal relationship can be established. Often the foundation for more definitive studies.
  • Case-control study, cohorts are defined by the presence or absence of the outcome, then prior exposures are identified and their odds compared (the mirror image of the cross-sectional design, which samples by exposure).
  • Case series, a report of a small group of patients sharing clinical features, generally without a control group. Prevalent in surgery (the Whipple procedure and the Nissen fundoplication both originated in case series) but weak because of selection, bias and confounding; useful for generating hypotheses for an RCT.
  • Expert opinion, the lowest level of evidence, representing individual experience and anecdote; before EBM it was the primary means of teaching medicine. It should be sought only in the complete absence of literature.

Thresholds and severity

Spending time refining the question is probably the most important part of the research process [1]. Once an idea is formed or a question asked, it must be transformed into a hypothesis, and it is worth asking whether the question posed really matters [1].

Length should follow the message. A paper should be as long as the size of its message: readers of large randomised multicentre trials need as much detail as possible, while reports of small simple trials should be brief [1].

NIHR Research Design Service · UK research collaboratives

The UK has a named free service for research design. The National Institute for Health Research runs the Research Design Service, which provides free and confidential advice on research design, writing funding applications and obtaining public engagement in research, for all researchers; training courses in research methodology and application are also available [1].

Trainee research collaboratives began in the UK and surgery led the way. The first trainee-level research collaborative was formed in 2008, when surgical trainees frustrated by the difficulty of conducting high-quality research during full-time training created the West Midlands Research Collaborative [1].

The premise solved two problems at once: to create and conduct prospective projects collating data across all members' units, and to take advantage of trainees rotating between units to ensure project longevity and enable longer-term outcome collection [1]. By achieving a critical mass of engaged members, collective momentum ensured completion even when individuals could not contribute consistently because of examinations, family life or busy clinical periods [1].

Where audit is concerned, UK practice is to seek advice from the local audit department and to ensure institutions have agreed to undertake the audit before it begins [1].

Snapshot audits are attractive because they cost almost nothing. Their key advantages are easy accessibility and the fact that they can be conducted at almost zero cost, making them an excellent way to bring a new group together and create contemporaneous real-world data [1].

Hierarchies of evidence: the pyramid, CEBM and GRADE

The original EBM architects drew a pyramid with expert opinion at the base and RCTs at the peak, in Schwartz's figure the tiers run, from bottom to top, expert experience/opinion, in-vitro research, animal research, case reports, case series, case-control study, cohort study, RCT [2]. The scheme was simplistic: it rested on the unproven assumption that RCTs are inherently superior to observational studies, it entangled the method of evidence collection with study design, and by placing systematic reviews above RCTs it failed to recognise that a systematic review can summarise any type of evidence, cohort studies, case-control studies, even case reports [2]. By 2002 more than 100 unique evidence-rating systems existed, and the differences are not trivial: in 2009 the American Association of Orthopedic Surgeons, working from the same data as the American College of Chest Physicians, could rate no VTE-prophylaxis recommendation for hip or knee surgery higher than grade B on level III evidence, whereas the ACCP gave it grade 1, level A [2].

  • The Oxford Centre for Evidence-Based Medicine (CEBM) Levels of Evidence (released 2000, updated 2011) is one of the most widely adopted systems.
  • It is organised as a table whose rows are clinical questions, How common is the problem? Is this diagnostic test accurate? What will happen if we do not add a therapy? Does this intervention help? What are the common harms? What are the rare harms? Is this early-detection test worthwhile?, and whose columns are the steps to search, from level 1 (strongest) on the left to level 5 (weakest) on the right [2].
  • For treatment benefit, level 1 is a systematic review of randomised or n-of-1 trials, level 2 a randomised trial or an observational study with dramatic effect, level 3 a non-randomised controlled cohort, level 4 case series, case-control or historically controlled studies, and level 5 mechanism-based reasoning; a level may be graded down for study quality, imprecision, indirectness or inconsistency, and graded up for a large effect size [2].
  • CEBM is meant as a pragmatic heuristic for answering questions in real time, allows resort to individual studies where other systems assume a systematic review exists, and covers prevalence, diagnostic accuracy, prognosis and screening as well as therapy and harm [2].
  • It is a hierarchy of the likely best evidence: an observational study with a large treatment effect may outweigh an inconclusive systematic review [2].

GRADE (Grading of Recommendations, Assessment, Development and Evaluation) classifies the quality of a body of evidence into four levels [2]:

GRADE qualityMeaning
HighFurther research is very unlikely to change confidence in the estimate of effect
ModerateFurther research is likely to have an important impact on confidence and may change the estimate
LowFurther research is very likely to have an important impact and is likely to change the estimate
Very lowAny estimate of effect is very uncertain
  • Quality is not assigned by study design alone.
  • An RCT starts "high" but may be demoted for study limitations, inconsistent results, indirectness, imprecision or reporting bias; an observational study starts "low" but may be upgraded for a large magnitude of effect, a dose-response relationship, or when all plausible biases would reduce the apparent effect, which is how GRADE allows observational data to establish causation where an RCT is unethical or unnecessary (alcohol and cirrhosis, asbestos and mesothelioma) [2].
  • GRADE then moves explicitly from evidence to recommendation through a summary-of-findings table presenting evidence quality alongside relative and absolute effects for patient-centred outcomes, a format designed to minimise framing effects, in which raters reach different conclusions from identical information presented as gain versus loss [2].
  • Recommendations are either strong (desirable effects clearly outweigh undesirable ones, or vice versa) or weak, and strength depends on more than evidence quality: uncertainty about the balance of benefit and harm, variability in patient values, and whether the intervention is a wise use of resources all count [2].
  • Hence a strong recommendation can rest on weak evidence, proton-pump inhibition for Zollinger-Ellison syndrome is strongly recommended because the benefit so clearly outweighs the risk, and a high-quality body of evidence for a small effect may yield only a weak one [2].
  • GRADE has been adopted by national and international societies, government health bodies, regulators and resources such as UpToDate, and its use has raised the rigour of systematic reviews; its costs are complexity and a steep learning curve [2].

Internal versus external validity is the caveat on every grade. Studies are judged on their internal validity (whether a causal conclusion is warranted for the study population) whereas applying a recommendation to a given patient depends on external validity, the generalisability of that conclusion beyond the original studies; all evidence must be applied within the context of the patient in front of you [2].

What a trustworthy guideline looks like

Schwartz lists nine qualities of the highest-quality, most clinically useful guidelines [2]:

  • 1.
  • An explicit, publicly available description of development and funding. 2.
  • A transparent process that minimises bias, distortion and conflicts of interest. 3.
  • A multidisciplinary panel of clinicians, methodologists and representatives (including a patient or consumer) of the populations affected. 4.
  • Rigorous systematic evidence review considering the quality, quantity and consistency of the aggregate evidence. 5.
  • A summary of potential benefits and harms for each recommendation. 6.
  • An explanation of the parts played by values, opinion, theory and clinical experience. 7.
  • A rating of confidence in the evidence and of the strength of each recommendation. 8.
  • Extensive external review including an open period for public comment. 9.
  • A mechanism for revision when new evidence appears.

Such guidelines are often interpreted as the standard of care, but several may apply to different aspects of one clinical situation, they cannot exist for every situation, and they must not be extrapolated beyond their specific conditions [2].

Treatment and Management

The audit cycle

Seven steps establish it [1]:

  • 1.
  • Define the audit question in a multidisciplinary team. 2.
  • Identify the body of evidence and current standards. 3.
  • Design the audit to measure performance against agreed standards based on strong evidence, seeking appropriate advice and institutional agreement. 4.
  • Measure over an agreed interval. 5.
  • Analyse the results and compare performance against the agreed standards. 6.
  • Undertake gap analysis, if all standards are reached, reaudit after an agreed interval; if improvement is needed, identify interventions such as training and agree them with the parties involved. 7. Reaudit.

Appraising a surgical randomised trial

Fewer than half of journal articles adequately report their study design, which led to the CONSORT (Consolidated Standards of Reporting Trials) guidelines in 1992, revised in 2010, a minimal set of reporting requirements (randomisation, blinding and so on) that many surgical journals now require as a completed checklist before an RCT manuscript is considered [2]. The 25-item checklist runs from identification of the study as randomised in the title, through eligibility, interventions, pre-specified outcomes, sample-size calculation and interim stopping rules, sequence generation, allocation concealment, who was blinded and how, participant flow with losses and exclusions, baseline data, numbers analysed and whether by original assignment, effect sizes with precision (both absolute and relative for binary outcomes), harms, limitations and generalisability, to registration number, protocol access and funding [2].

Internal validity (whether the results are accurate for the sample studied) has five components [2]:

  • Randomisation creates groups with similar known and unknown prognostic factors, eliminating selection bias (an investigator free to choose would assign the more favourable arm to particular patients). The method must be reported: quasi-random allocation by date of birth, day of the week or participant number is not truly random and cannot be concealed from study personnel. Even true randomisation balances confounders only in theory, a trial would have to be repeated indefinitely to guarantee it.
  • Blinding reduces performance bias, in which knowledge of allocation influences subjective outcomes (the placebo effect). Authors should state exactly which groups (subjects, clinicians, assessors) were blinded rather than write "double-blinded". Blinding is the major hurdle of surgical trials, given the ethical dilemmas of sham or placebo surgery, and is impossible when an operation is compared with non-operative management.
  • Equivalence among groups means each arm is treated identically apart from the intervention, the same number of visits, the same diagnostic tests.
  • Completeness of follow-up, attrition bias arises when withdrawals differ between groups, usually with an identifiable pattern (treatment, side effects, long follow-up), and those who remain may select for characteristics that determine efficacy.
  • Accuracy of analysis, analysing only completers skews conclusions, so most RCTs use intention-to-treat (ITT): "once randomised, always analysed." Removing non-compliers overestimates the effect size; since a proportion of real-world patients will be non-compliant, ITT better represents the population.

External validity (relevance and generalisability to the clinical population) is what actually changes practice [2]:

  • Number needed to treat (NNT) is the inverse of the risk difference between groups; the smaller, the more efficacious. In an RCT of laparoscopic cholecystectomy versus observation to prevent recurrent idiopathic acute pancreatitis, the NNT was five.
  • Number needed to harm (NNH), how many undergo the intervention for one adverse event; the higher, the safer. A low NNT with a high NNH is preferred, but neither captures the degree of benefit or harm.
  • Generalisability, restrictive inclusion criteria produce a homogeneous trial population whose results may not translate to heterogeneous "real-world" patients, and trials answer for the "average" patient when most patients are not average. Further studies in broader populations (the analogue of phase 4 pharmaceutical trials) and continued post-implementation data collection are needed before practice changes.

Logistical barriers are specific to surgery [2]:

  • Recruitment, accruing enough patients for adequate power is exponentially harder for rare diseases; expanding to multiple centres trades internal validity (more heterogeneity) for external validity.
  • Learning curves and expertise-based design, a drug is administered without measurable deviation, an operation is not. Novel procedures have learning curves even for experienced surgeons, and neglecting them underestimates the experimental intervention; established procedures carry each surgeon's individual habits. The expertise-based design keeps patients randomised but has each arm operated on only by surgeons expert in that procedure (already the norm in cross-specialty trials such as open versus radiological gastrostomy), at the cost of no longer modelling everyday practice.
  • All-or-none situations, a 2003 BMJ paper asked why there were no RCTs of parachutes in gravitational free-fall. Where everyone exposed to a risk suffers the outcome and no one treated does, an RCT is dangerous and unethical and observational data suffice.
  • Noninferiority trials, when an effective therapy exists, a placebo comparison is unethical, so trials aim to show a new therapy is not worse on the primary endpoint while improving secondary ones; the 2004 open-versus-laparoscopic colectomy trial for colon cancer sought equivalent oncological outcomes with better cosmesis, less pain and fewer hernias. Such trials grew from under 100 in 2005 to nearly 600 in 2015, and their critical parameter (the prespecified noninferiority margin) is largely arbitrary in the literature.

Procedural interventions

Writing it up

Articles are submitted in IMRAD form, introduction, methods, results and discussion [1]. Each section has its own discipline [1]:

  • Introduction, always short: a brief background, then the aims.
  • Methods, methodology and design in detail, identifying potential biases; new techniques described in full, established ones referenced rather than described.
  • Results, almost always best shown diagrammatically in tables and figures, and results shown in a diagram need not be duplicated in the text.
  • Discussion, do not repeat the introduction or reiterate the results; interpret the study intelligently and suggest future studies or changes in management. Do not indulge in flights of fantasy about future possibilities; most journal editors will delete them.
  • References, all relevant previous studies, not necessarily exhaustive but up to date, presented in the style of the journal being submitted to.

Publication

Most surgeons publish in peer-reviewed journals, where submitted work is checked anonymously by other surgeons before publication; editors will often advise on whether an article suits their journal [1].

Two funding models exist. It is usually free to publish in surgical journals, the cost of refereeing and editing being borne by the subscriber; open access, in which the author pays, ensures all research is visible to anyone by pushing editorial costs onto the study budget, and may become standard [1].

The statistics a reader must understand

Every significance test declares a null hypothesis (the "default" state, no difference) and can err in two directions [2]:

DecisionH₀ trueH₀ false
Reject H₀Type I error (false positive)Correct (true positive)
Fail to reject H₀Correct (true negative)Type II error (false negative)
  • The type I error rate, α, is the probability of rejecting a true null hypothesis, the significance level, conventionally 0.05.
  • The type II error rate, β, is the probability of failing to reject a false null hypothesis, and is tied to power (0 to 1): as power rises the chance of a type II error falls. Power depends on three factors (the significance criterion, the magnitude of the effect of interest, and the sample size) and power analysis gives the minimum sample size needed to detect an effect of a given size [2].

The P value (the probability of the observed result given that the null hypothesis is true) was Sir Ronald Fisher's innovation, and the "less than 1 in 20" threshold of P < 0.05 is arbitrary [2]. Schwartz's warnings [2]:

  • A P value is specific to that study's sample and may not generalise; the probability that a "positive" finding is actually false depends on the prior probability of the association and the study's power, not on P alone.
  • Statisticians estimate that a threshold of 0.05 leads to wrong conclusions at least 30% of the time, and more often in underpowered studies.
  • P values impose a binary verdict (is 0.049 meaningful and 0.051 not?) and say nothing about effect size, an intervention can be statistically significant and clinically irrelevant. Effect size, confidence interval and power must be read alongside P.
  • Despite these flaws their use keeps increasing; each reader should be sceptical and await replication. Fisher never endorsed the modern P < 0.05 criterion, he envisaged repeating experiments until the investigator could reliably reproduce the result.

Bayesian statistics are the main alternative, estimating from prior knowledge or data the probabilities that the hypothesis is true, that it is true given the observed data, that the alternative is true, and that the data would have been observed under the alternative, and combining them into a Bayes factor, the likelihood ratio of two competing hypotheses. Reliable priors are often unavailable, and a misinterpreted Bayes factor is as troublesome as a misinterpreted P value [2].

The kappa (κ) coefficient measures inter-rater agreement for qualitative items, correcting for chance: κ < 0 no agreement; 0–0.2 slight; 0.21–0.4 fair; 0.41–0.60 moderate; 0.61–0.80 substantial; 0.81–1 almost perfect [2].

Four questions every article should answer about its results [2]: What is the statistical significance? What is the effect size, and is it clinically relevant? What is the confidence interval? What was the power to detect a meaningful difference?

Complications

The failure modes named in the chapter are two: a project that fails because no team was built around it [1], and an audit that never closes its loop, since the reaudit is what converts measurement into improvement [1].

When evidence-based medicine fails its own standards

Schwartz's central paradox: there is no evidence that the EBM grading systems are themselves reliable [2].

  • External consistency, the GRADE Working Group compared six systems (ACCP, Australian NHMRC, Oxford CEBM, SIGN, US Preventive Services Task Force, US Task Force on Community Preventive Services) against 12 criteria and found poor agreement about the sensibility of the six systems; GRADE was built to overcome this but its ability to do so has never been formally assessed. The Surviving Sepsis Campaign illustrates the problem: rapid intravenous antibiotics were graded "E" (the lowest level, IV–V evidence) in 2004 and 1B/1D (a "strong" recommendation) in 2008, although the only three intervening studies were non-randomised and reached no different conclusions from earlier work [2].
  • Internal consistency, the 2005 GRADE pilot among 17 assessors found kappa values from 0 to 0.82, mean κ = 0.27 (fair), and κ < 0 for four judgements; the authors concluded that judgements about evidence are complex and could be resolved by discussion, and no assessment of GRADE's reliability has been published since [2].
  • System issues, GRADE deliberately separates evidence quality from recommendation strength, which permits "high-quality" evidence for trivial effects and strong recommendations on low-quality evidence; and the up- and down-grading factors (blinding, follow-up, consistency, generalisability, effect size) are fundamentally different quantities that cannot simply be added or subtracted, so weighting them is individual judgement [2].
  • Validity, none of the GRADE publications provide validation or proof of usefulness; no RCT of the effect of EBM on patient outcomes has ever been undertaken, so EBM does not satisfy its own requirements and is, in Schwartz's words, "a form of systematic expert opinion" [2].
  • Unintended consequences of strong recommendations, they can foreclose debate and research on a misguided topic. Antibiotic prophylaxis in necrotising pancreatitis was strongly recommended on the basis of multiple RCTs, meta-analyses and systematic reviews, then reversed when further trials showed no benefit; strong recommendations can render life-saving prospective studies "unethical", and some groups have warned against converting guidelines into law [2].
  • The crisis of reproducibility.
  • Index publications underpinning fundamental precepts of practice or whole directions of drug discovery have proved impossible to reproduce; estimates of irreproducibility range from 75% to 90% by mathematical inference, and in one practical investigation 0 of 52 observational findings were confirmed by RCTs [2].
  • As bias increases, the positive predictive value of a finding falls; medical research operates in areas of low pre- and post-study probability, so observed effects varying around the null may simply measure the prevailing bias of a field.
  • With many teams pursuing the same question worldwide, the chance that at least one claims significance rises while the PPV falls, and the first to report receives disproportionate attention; under the current framework a PPV above 50% is hard to achieve, and even a well-constructed, adequately powered RCT with a 50% pretest probability reaches a true conclusion only about 85% of the time [2].

Medical reversal, a term introduced by Vinay Prasad and Adam Cifu in 2011, describes an established practice or drug falling out of favour because it is subsequently shown not to work; its central EBM issues are surrogate endpoints, misrepresentation of trial effects, academic and economic bias in reporting and dissemination, and the reliability of alternatives to the RCT [2]. Schwartz notes that most of these problems trace straight back to Bacon's idols of 1620 [2].

Outcomes

Snapshot audits generate hypotheses and identify areas where further prospective research is needed, by exploring differences in patients, techniques and management across a cohort to find practice variability that may explain differences in outcome [1].

The thrombolysis example is the argument for the whole enterprise: the evidence existed and the synthesis did not, and patients died in the interval [1].

Schwartz's closing position: guidelines as rules of thumb

EBM is appealing because it reduces uncertainty, but its performance in improving patient care has never been validated, its systems are inconsistent, and a single system can assign varying grades on subjective factors [2]. As medicine moves toward "precision" and "individualised" care, EBM's focus on guidelines for the average patient becomes a paradox; the best physicians work from scientific theory expressed through practical, locally acquired tacit knowledge, and "EBM itself is not founded in scientific principle" [2].

The alternative Schwartz proposes is common-sense application of scientific principles with healthy scepticism: use published guidelines within the context of one's own practice until a grading system is shown to improve outcomes, recognising that daily practice involves hundreds of decisions mixing explicit and tacit knowledge, and that guidelines should be flexible and receptive to feedback from practising clinicians [2].

What researchers can do: obtain better-powered evidence (high-powered, low-bias meta-analyses approach a theoretical gold standard, though even they carry bias and large-scale evidence is not always possible); stop judging any single study in isolation and instead weigh the entire body of evidence, ideally by networking data across research groups, a significant change in academic culture [2].

What surgeons should do: ask "What is the best course of action for this patient, in these circumstances, at this point in their illness?", synthesise the evidence and then individualise it through the ethics, personality and values of the case. Risk calculators inform discussion but are not definitive evidence for or against a treatment; judgement remains necessary, and guidelines should be thought of as "rules of thumb" that require context rather than "rules of law" [2].

References

  1. Bailey & Love's Short Practice of Surgery, 28th ed., Ch. 13 Surgical audit and research
  2. Schwartz's Principles of Surgery, 11th ed., Ch. 51, Understanding, Evaluating, and Using Evidence for Surgical Practice