top of page

Metrology, Calibration and Test Statistics

Aug 20
12 min read

Any voltmeter shows a value. A laboratory declares a result.

This is the module that separates laboratories that issue numbers from laboratories that issue reports. Any voltmeter indicates a value. A competent laboratory declares a result, with traceability, uncertainty and a decision rule attached. Modules 3 through 8 covered generating and measuring high voltage and current. Module 9 covered protecting those measurements from the parasitic phenomena of the installation. What is missing is the step that turns an instrument reading into something defensible in front of a customer, an accreditation assessor, or a court.

The answer has three layers. Traceability: the indicated value must connect, through a documented and unbroken chain of calibrations, to the realization of the volt in the SI. Uncertainty: no measurement is exact, and IEC 60060-2 imposes quantitative limits — typically 3% on voltage for an approved system — that can only be demonstrated with a budget built to the GUM. And statistics: the breakdown of insulation is a stochastic phenomenon, so the breakdown voltage of an object is not a number but a probability distribution, and testing means estimating the parameters of that distribution from a finite number of applications. By the end, the challenge set at the end of Module 9 is closed: two laboratories can publish two different U50 values for the same object and both be correct.

What you will be able to do

  • Distinguish an approved measuring system (AMS) from a reference measuring system (RMS) per IEC 60060-2:2025, and explain why approval is tied to voltage shape, polarity, range, physical configuration and software version — not to the divider alone.

  • Perform a comparison calibration of an AMS against an RMS: choose the voltage levels with h ≥ 5 across the full range, compute the assigned scale factor, and evaluate linearity by direct and alternative methods.

  • Build a test uncertainty budget to the GUM (JCGM 100:2008): Type A and Type B components, normal, rectangular and triangular distributions, combined and expanded uncertainty with k = 2.

  • Apply statistics to disruptive tests: breakdown probability function, normal, Weibull and Gumbel distributions, the characteristic voltages U50, U10 and Uw, and the 15/2 acceptance procedure.

  • Run and process an Up & Down series, apply the Dixon-Mood estimators for U50 and sigma, and judge the statistical quality of the series.

  • Structure an auditable report to ISO/IEC 17025:2017: mandatory content, decision rule, evidence and amendment control.

Traceability is a chain, not a sticker

The VIM defines metrological traceability as the property of a result by which it can be related to a reference, generally the realization of an SI unit, through an unbroken and documented chain of calibrations, each contributing to the measurement uncertainty. That definition is demanding and it excludes the everyday industrial shortcuts. Traceability is not owning a certificate, not using a well-known brand of instrument, not displaying an in-date calibration label, and not holding an ISO 9001 certificate. It is the chain, with its uncertainties added link by link, that has to be technically demonstrable.

At the top sits the quantum realization of the volt through the Josephson effect, with DC voltage standards at uncertainties of parts in 10⁹. Below it, national metrology institutes hold the national link for high voltage; below them, calibration laboratories accredited to ISO/IEC 17025 transfer traceability to a testing laboratory's reference system, which in turn calibrates the working systems. Accept a certificate only after checking that the laboratory was accredited on the calibration date, that the quantity and range are in the published scope, that the declared uncertainty and best measurement capability suit the intended use, that the calibrated configuration matches the one used in service down to cable and software version, and that any correction to apply and any restriction on polarity, frequency or waveform is stated. Calibration intervals are set by risk, stability, severity of use and the consequence of a wrong measurement, not by a universal rule; shorten them after an overload, a flashover near the divider, a rough shipment, a repair, or any change of cable, instrument or software.

Approved and reference systems: approval is never universal

A high voltage measuring system, in the language of the standard, is the whole chain: converting device (resistive, capacitive, mixed or universal divider, voltage transformer, measuring capacitor), transmission system (coaxial cable with its impedance and termination, attenuators, adapters, sometimes fiber), recording instrument to IEC 61083, and evaluation software with its peak, T1, T2, overshoot and k(f) algorithms. The overall scale factor is the product of the partial ones, F_system = F_converter · F_transmission · F_instrument, and the measured value is U = F_system · X.

An approved measuring system is one that has demonstrated compliance through performance tests documented in its performance record, within declared conditions and ranges. The point that fails candidates in audit is that approval is not universal: it is tied to voltage shape, polarity, amplitude range, frequency or time parameters, physical configuration including height and proximity, operating time, maximum application rate, the specific cable and channel, and the software version. A divider approved for positive lightning impulse from 200 kV to 1 MV is not automatically approved for negative impulse, for front-chopped impulse, for AC, for another cable or for another digitizer. A reference measuring system, by contrast, carries traceable calibration, uncertainty small enough to calibrate an AMS by comparison, demonstrated stability and controlled configuration. The relationship is asymmetric: a reference system may serve as an approved system in ordinary testing, but an approved system never becomes a reference system by decree.

The uncertainty limits are the table that governs the laboratory. Expanded uncertainty (k = 2) is limited to 3% for DC mean value, AC, full or tail-chopped lightning impulse peak and switching impulse peak on an approved system, against 1% for a reference system; 5% for front-chopped lightning impulse peak, against 3%. Time parameters carried a 10% approved-system baseline, and the change with the greatest practical impact in IEC 60060-2:2025 is that the limit for the front time T1 of the standard lightning impulse rose from 10% to 15%, an acknowledgement of how hard T1 really is to measure with the virtual-front definition on systems of finite dynamic response. Two traps: the limits apply to the uncertainty of the measurement in use, not to the calibration certificate alone, and the common claim that the standard requires 5% on time parameters confuses the reference system requirement with the approved system requirement.

Calibration by comparison, and the myth of ten readings

The preferred calibration is simultaneous comparison against a reference system: both measure the same applied voltage at the same time. For each application i at level g, the individual factor is F_i,g = U_ref,i,g / X_AMS,i,g. The layout must suppress the systematic error sources of Modules 7 and 9 — unequal electromagnetic coupling, inductive drops in connections, ground loops, cable reflections, mutual proximity influence between the two dividers, corona, and differences in bandwidth or acquisition instant between channels.

Get the nomenclature right, because one classic procedural error lives here. h is the number of voltage levels, g is the level index running from 1 to h, and n is the number of independent determinations at each level. Full-range calibration requires h ≥ 5, including the lower limit, the upper limit and at least three intermediate levels. At each level you make n independent determinations, and the standard's note that more than ten are normally unnecessary is a usual ceiling, not a mandatory minimum. So h = 5 is the minimum number of levels, while "g = 10 readings" is a confusion of symbols.

At each level compute the mean factor F_g, the relative experimental standard deviation s_r,g, and the Type A standard uncertainty of the mean, u_A,g = s_r,g/√n. The assigned factor for the range is F = (1/h)·ΣF_g, and a conservative approach takes the largest u_A,g as the Type A component. Non-linearity is treated as a rectangular distribution over the largest relative deviation, u_nl = (1/√3)·max|F_g/F − 1|. Worked example on a nominal 10,000:1 divider calibrated at five levels reading 10,004 / 10,009 / 10,011 / 10,016 / 10,022: F = 10,012.4, the largest relative deviation is 0.0959%, and u_nl = 0.0959%/√3 = 0.0554%. Report the individual level factors, not just the mean — the mean alone hides exactly the information the user needs to build a use-case budget. When the reference system does not reach the full range, combine a levels of direct comparison with b levels of linearity testing under a ≥ 2, b ≥ 2 and a + b ≥ 6, with direct comparison reaching a representative voltage.

The GUM budget that supports the 3%

The most common error in applying the GUM is starting with the spreadsheet. Define the measurand and the measurement model first. For voltage measured by a divider: U_H = F · U_L · C_T · C_P · C_D · C_S, where F is the assigned scale factor, U_L the low-voltage indication, and C_T, C_P, C_D and C_S the corrections for temperature, proximity, drift and software processing. Taking a correction as unity, which is usual, does not remove its uncertainty from the budget.

Type A evaluation uses statistics of repeated observations, u_A = s/√n. Type B uses any other information: certificates, manufacturer specifications, resolution, history, environmental limits. The classification describes the evaluation method, not the physical nature of the error — a systematic effect can be evaluated by a Type A method. Converting information to standard uncertainty: a certificate stating U and k gives u = U/k; symmetric limits ±a with equally likely values give u = a/√3; limits with central values more likely give u = a/√6; a digital resolution d gives u = d/√12. Combine as u_c²(y) = Σ c_i²·u²(x_i), and do not ignore correlation when two instruments share a calibration standard or one time base affects both T1 and T2.

A worked case at 800 kV, with F = 10,000 and U_L = 80.000 V. Components in relative terms: factor calibration 0.3000%, repeatability 0.0316%, digitizer calibration 0.0500%, resolution 0.0036%, factor drift 0.1155%, temperature 0.0866%, proximity 0.0577%, non-linearity 0.0816%. The quadratic sum is u_c,r = 0.3526%, so u_c = 2.821 kV and, with k = 2, U = 5.6 kV or 0.71% relative. Declare it as U_H = (800.0 ± 5.6) kV, k = 2, coverage probability approximately 95%, comfortably inside the 3% limit. Read the hierarchy: factor calibration dominates at 0.30%, followed by drift at 0.12%. Buying a better digitizer, already at 0.05%, without attacking divider drift would be optimizing in the wrong place — a direct lesson for the laboratory's investment plan.

Breakdown is a distribution, not a number

Self-restoring insulation does not break down at the same voltage every time. The process depends on the availability of initial electrons, the development of avalanches and streamers, surface condition, particles and humidity, polarity and waveshape, the interval between applications, and the dielectric's memory of space charge from earlier discharges. A single breakdown voltage therefore does not characterize the object. What characterizes it is the breakdown probability function Pd(U), and the test exists to estimate its parameters.

Under the normal model, Pd(U) = Φ[(U − U50)/σ]. U50 is the 50% breakdown voltage, U10 the 10% breakdown or 90% withstand point, and σ the dispersion. Since the 10% quantile of the normal distribution is z = −1.28155, U10 = U50 − 1.28155·σ exactly, and insulation coordination practice uses the approximation Uw ≈ U50 − 1.3·σ for the statistical withstand voltage. With U50 = 1,000 kV and σ = 30 kV, that is 961.55 kV exact against 961 kV approximate — a 0.06% difference. The normal is not the only candidate. A three-parameter Weibull, Pd(U) = 1 − exp{−[(U − γ)/η]^β} for U > γ, is preferable for weakest-link behaviour such as large insulating areas or cables, or when the asymmetric lower tail matters; the Gumbel distribution models extreme values. Do not choose on best R² alone: the physical mechanism, the observed probability range, sample size, censoring and independence between applications all weigh in.

The most widespread impulse withstand test is the 15/2: apply 15 impulses at the specified level and accept self-restoring insulation with at most 2 breakdowns and no failure in any non-self-restoring part. Understand what it is — acceptance at one level — and what it is not — a method of estimating U50 and σ. If the true per-application breakdown probability is q, the probability of acceptance is the cumulative binomial over j = 0 to 2 of C(15,j)·q^j·(1−q)^(15−j). That gives 96.4% acceptance at q = 5%, 81.6% at q = 10%, 39.8% at q = 20% and 12.7% at q = 30%. An object with a genuine 10% breakdown probability per impulse passes more than four times out of five, so passing bounds the probability statistically but never proves it is zero.

Up & Down and the Dixon-Mood estimators

IEC 60060-1 organizes disruptive test procedures into multiple-level, up-and-down and progressive stress. The Up & Down method is the most efficient way to estimate U50 because it automatically concentrates applications around U50 itself. Estimate U50 in advance and start near it, fix a constant step ΔU, apply one impulse, drop by ΔU after a breakdown and raise by ΔU after a withstand, hold polarity, waveshape, interval and ambient conditions constant, and continue to the required number of valid applications — at least 20 in the usual single-application grouping. Record every application, including the invalid ones, with justification.

Step size matters more than the textbooks suggest. The commonly repeated "3% to 5% of U50" is not a universal rule; the literature also records 1% to 3% and 1.5% to 3%. The correct principle is to size the step against the expected dispersion of the object, with 0.5·σ ≤ ΔU ≤ σ as the practical guide. Too small a step confines the series to very few levels, lengthens the test and degrades the σ estimate. Too large a step coarsely quantizes the series into artificial alternation between two levels, biases U50 and makes σ dominated by the step itself.

With equally spaced levels U_i = U_0 + i·ΔU, count at each level the number n_i of occurrences of the less frequent event of the series, then form N = Σn_i, A = Σi·n_i and B = Σi²·n_i. Then U50 = U_0 + ΔU·(A/N − ½) when the counted event is breakdown, or with +½ when it is withstand, and σ = 1.620·ΔU·[(N·B − A²)/N² + 0.029], valid only under the quality condition (N·B − A²)/N² ≥ 0.3. Worked example: 20 valid applications with ΔU = 30 kV over levels 960, 990, 1,020 and 1,050 kV, with 9 breakdowns distributed 0, 2, 5 and 2. Then N = 9, A = 18, B = 40, U50 = 960 + 30·(2 − 0.5) = 1,005 kV, the quality index is 36/81 = 0.444 which passes, σ = 1.620·30·(0.444 + 0.029) = 23.0 kV or 2.3% of U50, U10 = 975.5 kV and Uw ≈ 975.1 kV. Rule of thumb for choosing the method: Up & Down is unbeatable for economy of impulses when estimating U50, but to defend a U10 or a low-probability withstand in front of a third party, multiple-level testing with ten or more impulses per level and maximum-likelihood fitting produces a statistically stronger argument.

Syllabus

  • Metrological traceability as defined in the VIM, and the shortcuts that are not traceability

  • The chain from the SI realization of the volt through national metrology institutes and accredited calibration laboratories

  • Accepting a calibration certificate: the ten questions that belong in the quality procedure

  • Calibration intervals set by risk, not by calendar; intermediate checks and their acceptance limits

  • The measuring system as converting device, transmission system, recording instrument and evaluation software

  • Approved measuring system: the performance record and why approval is never universal

  • Reference measuring system: traceability, low uncertainty, demonstrated stability, controlled configuration

  • The uncertainty limit table: 3% on voltage, 5% on front-chopped peak, and the 2025 change of T1 from 10% to 15%

  • Assigned scale factor, assigned ranges and assigned operating conditions

  • Comparison calibration: h ≥ 5 levels, n independent determinations, and the h/g/n nomenclature

  • Level factors, relative experimental standard deviation, Type A uncertainty and the non-linearity component

  • Limited-range calibration: a + b ≥ 6 and the five accepted linearity methods

  • The GUM: measurement model first, then Type A and Type B evaluation and distribution divisors

  • Combined uncertainty, sensitivity coefficients, correlation, and expanded uncertainty with k = 2

  • A complete worked budget for a measurement at 800 kV

  • The stochastic nature of breakdown; the probability function Pd(U); U50, U10, U90 and sigma

  • Weibull and Gumbel distributions and how to choose one

  • The 15/2 procedure as acceptance, not estimation, and its binomial characteristic

  • Up & Down: procedure, step size selection, Dixon-Mood estimators and the quality condition

  • The auditable report: ISO/IEC 17025 clause 7.8 content, decision rules and guard bands

Laboratory work

An integrating exercise in three stages, run on data. Stage 1, calibration: you receive simultaneous readings from a reference system (F_ref = 1,000.0, U = 0.8%, k = 2) and from the divider under calibration at five levels from 200 to 1,000 kV with eight determinations per level, and you compute the level factors F_g, the assigned factor F, the relative standard deviations s_r,g, the Type A component u_A and the non-linearity component u_nl. Stage 2, uncertainty budget: using the simulated certificate and the stated use conditions — drift ±0.2%, temperature ±0.15%, proximity ±0.1%, digitizer U = 0.1% — build the complete GUM worksheet, declare U with k = 2, and verify compliance with the 3% limit of IEC 60060-2:2025. Stage 3, Up & Down: process a real 25-impulse series recorded as a withstand and breakdown sequence with levels, build the Dixon-Mood counting table, compute U50 and σ, verify the (N·B − A²)/N² ≥ 0.3 condition, derive U10 and Uw, and write the result paragraph of the report with every mandatory declaration: identification of the approved and reference systems with certificate, date and scope; assigned scale factor and range; circuit diagram with distances and earthing point; cable, termination, channel, sampling rate and bandwidth; software version, algorithm and k(f) parameters; atmospheric conditions and correction factors with the calculation method; impulse-by-impulse data; and the full statistics including exclusion criteria.

Key takeaway

Two laboratories test identical objects and publish U50 = 1,005 kV and U50 = 985 kV. Who is wrong? Possibly nobody, and there are at least five legitimate explanations, all quantifiable. Measurement uncertainty: if each laboratory runs at U = 2% (k = 2), that is ±20 kV at 1,000 kV, the declared intervals overlap widely and the compatibility index |y1 − y2|/√(U1² + U2²) = 20/28.3 = 0.71 is below 1. Statistical uncertainty of the estimator: with 20 impulses and σ = 30 kV, the standard error of U50 from an Up & Down series is on the order of 7 to 10 kV, and it adds to the budget. Atmospheric corrections applied with different density and humidity data. Conditioning and memory of the object, since different application sequences leave different surface states and space charge. And different step sizes and exclusion criteria in the two procedures. The maturity lesson: results can only be compared together with their uncertainties, their methods and their conditions, and a report that declares all three is exactly what lets an engineer judge compatibility instead of arguing about numbers.

Module 10 of the Atlas Energy Academy · The Lab as a System · about 6 hours · Video series plus numerical exercises · Prerequisites: Module 9 — Parasitic Phenomena and Test Installations

Recent Posts

See All

Comments


bottom of page