Skip to main content

Every NDT procedure rests on an assumption that is almost never tested: that the technique would find the defect if the defect were there. A clean report proves the technician followed the procedure. It does not prove the procedure works.

Flawed specimens are how that assumption gets checked. They are components containing known defects, of known type, size and position, made deliberately — and they are the only way to demonstrate that an inspection detects what it is supposed to detect.

The problem with a clean result

Consider a phased array procedure written for a thick-section narrow-gap weld. The scan plan looks right. The calibration block gives good responses. The technician is certified, the equipment is in date, the report says no recordable indications.

Every one of those things can be true while the focal laws leave a band of the fusion face uninterrogated. The scan produces a confident, complete-looking image of the wrong volume. Nothing in the process catches it, because everything in the process is checking that the procedure was followed rather than that it works.

Run the same procedure over a specimen containing a known lack of side-wall fusion at a known depth, and you find out in an afternoon.

That is the whole argument. Flawed specimens turn "we followed the procedure" into "the procedure finds the defect".

Why a drilled hole is not a defect

The obvious way to make a flawed specimen is to machine something into a plate. Side-drilled holes, flat-bottomed holes and machined notches all have their place — they are reproducible, cheap and precisely locatable, which makes them excellent for calibration.

They are poor at representing real defects, because real defects do not behave like machined features.

A side-drilled hole is a smooth cylinder. It reflects ultrasound predictably from every angle. A genuine lack of fusion is a rough, tight, planar discontinuity that may be partially closed, sits at whatever angle the joint preparation dictated, and reflects strongly at one angle and barely at all at another. A procedure qualified against drilled holes can pass comfortably and still miss the thing it was written for.

The same is true of cracks. A machined notch is an open slot with parallel sides. A real fatigue crack is tight, branched, and its faces touch — which lets ultrasound cross it rather than reflect from it. Notches consistently over-estimate detectability.

So the useful distinction is between calibration targets, which are meant to be reproducible, and representative flaws, which are meant to behave like the real thing. Procedure qualification needs the second.

How representative flaws are actually made

Several techniques, chosen by what the flaw needs to imitate.

Welded-in defects. The most common route for weld inspection specimens. The defect is produced by deliberately mis-welding — dropping heat input to cause lack of fusion, adjusting the root gap to leave lack of penetration, introducing contamination to generate porosity, or leaving slag between passes. Done by a skilled welder working to a deliberately wrong procedure, these are genuine welding defects, because they were made the way welding defects are made.

The difficulty is control. Mis-welding reliably enough to place a 12 mm lack of fusion at a specific depth, and no other defects anywhere else, takes practice and a lot of rejected attempts.

Implanted flaws. A defect is created separately — often by controlled fatigue cycling to grow a real crack — then a block containing it is inserted into a component and welded in, with the surrounding weld made correctly. This gives a genuine tight crack at a precisely known position. More expensive, and the implant boundary itself must not become a detectable feature.

Grown fatigue cracks. A starter notch is machined, the component is cyclically loaded until a crack initiates and grows from it, and the notch is then machined away. What remains is a real fatigue crack with real crack faces and a real tip. This is the gold standard for anything qualifying detection of service-induced cracking, and it is the slowest and most costly method.

Thermal fatigue and stress corrosion cracking under controlled conditions, where the qualification specifically concerns those mechanisms and their characteristic branching morphology matters.

The flaw map is the product

A specimen without documentation is a lump of metal with a secret.

Every specimen is supplied with a flaw map recording each defect's type, position in three axes, through-wall extent, length, orientation, and how it was created — plus the parent material, weld procedure and any heat treatment. That record is what makes the specimen usable as evidence, and it is what an auditor will ask for.

The map is verified independently. A defect intended to be 8 mm through-wall is confirmed by an independent method — often radiography during manufacture, and ultimately by sectioning a sister specimen and measuring it under a microscope. Destructive verification of a representative sample is what converts an intention into a measurement.

This is where destructive testing and specimen manufacture meet, and why they sit in the same facility rather than being bought separately. The metallography that proves the flaw is what the flaw map claims uses the same equipment as the macro and micro examination on a welder qualification.

What specimens get used for

Procedure qualification. Demonstrating that a written technique detects the defect types and sizes it was written to find, at the sensitivities claimed. Increasingly required rather than assumed, particularly in nuclear.

Operator qualification and performance demonstration. Testing the person rather than the paperwork. A technician runs the qualified procedure over specimens whose contents they do not know, and their results are scored against the flaw map. This finds things a written examination never will.

Training. The difference between a trainee who has read about lack of side-wall fusion and one who has found it, sized it and been told they were 2 mm out, is the difference between knowledge and competence.

Equipment and technique comparison. Assessing whether a new instrument, probe or technique genuinely outperforms the incumbent, on the same specimens, with a known answer.

Round-robin and inter-laboratory trials. Where several parties inspect the same specimens and results are compared — which is uncomfortable and extremely informative.

Looking after them

Flawed specimens are expensive, long-lived assets, and they are routinely treated like offcuts.

They corrode. A carbon steel specimen left in a damp store develops surface rust that changes the couplant contact and, over time, the surface condition the technique was qualified against. Specimens need storing dry, lightly protected, and inspected periodically for their own condition.

They get damaged. Dropped, scored by careless handling, or ground by somebody wanting a better couplant surface — which alters the very geometry the flaw map describes.

They leak. Not physically. Once a specimen has been used for enough operator assessments, its contents become known — informally, between technicians, entirely without malice. A specimen whose flaws everyone knows is a training aid, not an assessment. Sets used for qualification need rotating, and the rotation needs managing by somebody who is not being assessed.

They need traceability. A specimen without its flaw map, or with a map nobody can match to a serial number, is worthless. Marking should be permanent, on the specimen, and the register should be a controlled document.

Organisations that build a specimen set and then let it decay tend to conclude that specimens were not worth the money. Usually what was not worth the money was buying them and then not owning them.

Specifying a set

The specimens that get used are the ones designed around a question. Worth settling before manufacture:

  • Which defect types matter for this application, and which are you willing to accept you will not find?
  • What sizes — including deliberately marginal ones. A set where every defect is comfortably detectable proves very little.
  • What material, thickness and joint geometry, matching the production component rather than a convenient plate.
  • How many blanks — specimens containing nothing. Without them, a performance demonstration only measures whether someone can find defects, not whether they call defects that are not there.

That last one gets left out most often, and it is the one that turns a training aid into an assessment.

Where we come in

Flawed, within the Test Inspect Group, designs and manufactures bespoke flawed specimens — weld joints, plates and components — for procedure validation, operator assessment and training, with controlled defect placement and a verified flaw map for each.

The destructive testing and metallurgical analysis that verifies those flaws is done in the same laboratory, which is the practical reason the two capabilities sit together: the specimen and the evidence that it is what we say it is come from one quality system.

If you have a procedure you have never actually tested against a known defect, that is the conversation worth having — and it usually starts with which defect would hurt most if you missed it.


TECHNICAL REVIEW — DELETE FROM THE LINE ABOVE, DOWN

Scheduled for 2026-09-03. Not to go live until a Level 3 or the RPA has read it.

Check: FLAWED sub-brand piece. Almost no comparable UK content exists, so this is a genuine differentiator. Confirm which manufacturing routes we actually offer (welded-in / implanted / grown fatigue cracks) before the article implies all three.

Image to shoot: Specimen set laid out with its verified flaw map alongside — the pairing is the point.

Internal links already in the text: /flawed/, /flawed-specimen-design-manufacture/, /inspection-validation-qualification/, /macroscopic-and-microscopic-examination/ — confirm they read naturally, do not add more.

Leave a Reply