Beyond the Instructions: Why the Best STEM Kits Miss the Mark

Thu, Oct 01, 2026 at 1:20PM

Beyond the Instructions: Why the Best STEM Kits Miss the Mark

Somewhere in your building there’s a cabinet with four or five kits in it that got used twice. Molded trays, a cutout for every part, laminated cards with eight numbered steps and a QR code. Nothing defective about any of them. They worked exactly as advertised, and that’s the problem.

The instinct to buy that kind of kit makes sense. You’ve got forty-five minutes, twenty-six students, and a curriculum that doesn’t care whether the printer jammed. Something that produces clean data inside one period and gets packed away before the bell is an easy solution.

Then you ask the class the following week what they learned, and not much comes back.

Most STEM kits are built to prevent confusion, but confusion is where learning happens.

Six parts in a bag and no diagram

Order the three-point flexure setup for the MSET and you get two quarter-inch thumbscrews half an inch long, a third one at three-quarters, a load cell extension, two supports, a load nose, and a strip of aluminum about the size of a paint stirrer. The parts are lettered A through F. Nothing in the bag says which piece goes where.

A student holding that load nose for the first time has no idea what it does.

That minute she spends not knowing is worth more than the forty that follow it. She’ll have to find the way out on her own, once she stops waiting for a diagram.

Kapur has spent twenty years working on this

Manu Kapur is a learning scientist at ETH Zurich. Since the mid-2000s he’s been running the same swap over and over, which is either teach the concept and then assign the problem, or assign the problem first and teach afterward.

Two controlled studies, published in Cognitive Science in 2014. Students who were given the problem first came out level with everyone else on routine calculation, and ahead on conceptual questions. They were faster on transfer problems, the kind nobody had shown them how to solve.

They also said it was harder, and that they put in more effort. Same students, same week, both results sitting there next to each other, which is inconvenient if you happen to be the person reading end-of-term surveys in June.

Kapur and Tanmay Sinha later gathered 53 experimental comparisons of the swap. It held. Got stronger, in fact, when the students struggled.

They’ll tell you the confusing kit was the worst one

There’s a companion study out of Harvard that explains why hardly anybody teaches this way. In 2019, Louis Deslauriers and his colleagues split an introductory physics course, running one half as active problem solving and the other as lecture with an experienced instructor working from the same material. The active students learned more. They also rated their own learning lower. Effort registered with them as poor teaching, so poor teaching is what went on the form.

Expect the same from your class. They’ll mean it, too.

Worth knowing before somebody with a clipboard walks in, as well. Four groups stuck and a teacher pointedly not helping does not read as good instruction from the doorway. Have that conversation with your principal in September.

If the argument you need is for a budget meeting instead, Scott Freeman’s group at the University of Washington pooled 225 studies of undergraduate STEM courses. Failure rates ran 33.8 percent under lecture and 21.8 percent where students worked actively. Exam scores about six percent higher. Half a letter grade for talking less, which decides more purchase orders than anybody will admit on a call.

The support span is the experiment

Somebody has to decide how far apart those two supports go, and if the fixture shows up assembled, that somebody was a technician in Warner, New Hampshire, before the box ever shipped.

Bad deal for the student. The distance between the supports is most of what the flexure test is about. Set them close and the beam barely moves; spread them out and the same piece of aluminum goes soft on you. That surprises students, because they’ve been taught that materials have properties and nobody ever mentioned that geometry gets a vote.

She learns it by sliding a support two inches and watching the number change. An assembled fixture hands her a span she didn’t pick, won’t question and has no reason to think about, and the data comes out cleaner for it. No catalog calls that a trade.

Swap her sample after the first run. Thinner aluminum, or something else at the same thickness, and have her write down what she thinks will happen before she loads it. On paper, in pen. Spoken predictions vanish the second they turn out wrong, and written ones sit there looking at you. The flexure procedure states its objective plainly enough: look at how beams of different shapes and materials behave across a support span, and see what the span width does. That’s a question rather than a recipe, and a harder document to write than eight numbered steps.

The case for buying one machine instead of nine

The MSET is a benchtop trainer. One frame, one load cell, and what changes from lab to lab is whatever gets bolted to it. Buckling, tension, spring stiffness, dynamic impact, cantilever and three-point flexure. Friction, magnetic force, magnetic damping, density. Pendulum dynamics and simple harmonic motion with mass loading. Buoyancy, hydrostatic pressure, lever balancing, load cell calibration.

Run buckling early if the sequence is yours to set. A slender column takes load, takes load, takes load, and then goes over sideways all at once, faster than anybody in the room is ready for. Students who have only met failure in a textbook are expecting something slow and sad. No diagram prepares them for the speed of it.

The bigger payoff is dull to describe, which is probably why nobody sells on it. By March, when the assignment is buoyancy, she’s working with the load cell she calibrated back in September. She knows how it behaves. She knows it reads funny mounted a hair crooked. There’s no fresh instruction sheet to hide behind, because the machine is the same machine and only the question has moved.

Whether a student can carry September into March is what a hiring manager digs at in a first interview, though nobody in the room calls it transfer.

The same unit gets ordered by middle schools and by universities, with the documentation written to the level.

Where this argument gets borrowed by people selling bad equipment

Difficulty is worth defending when the physics is causing it. A part that never made it into the box is a shipping error. Nobody learns anything from a shipping error, and software that won’t install on eight-year-old district laptops isn’t productive struggle either.

Kapur is blunt about the more serious limit. Struggle with nothing behind it doesn’t work. The back half of the design, where you explain what happened and why, is where it consolidates, so if the bell rings on twenty minutes of floundering and that’s the end of it, you’ve spent a period convincing a fourteen-year-old that she isn’t cut out for science. She’ll believe you.

Which is why the paperwork deserves as much attention as the hardware. Every MSET experiment ships with an instructor manual alongside the student procedure, covering background, setup, data acquisition and analysis. You know where the period is going. She doesn’t, and that gap is the design.

One more item, and it’s the least impressive line on any evaluation sheet. Find out whether the vendor sells replacement parts, and whether they’ll repair a unit that came back from a period with twenty-six ninth graders in it. Mentis does both, and there’s a form for it. Anybody planning for breakage has assumed students will put hands on the equipment.

Don’t fix anything for the first fifteen minutes

Hands go up early. Most of those questions don’t deserve a straight answer, so ask what they expected to happen, or what changed between the first run and this one, and then go stand somewhere else in the room.

A group that finishes fast doesn’t need a different activity. They need a shorter span or a thinner sample or a prediction they have to defend out loud.

Budget the time honestly, though. Twenty minutes of floundering needs the back half of the period for the explanation, or the next period. If you don’t have it, run something else that day.

Grading

If the point of the exercise was that the first three attempts failed, a rubric built around the correct final number cancels everything the lab was for. Students work that out inside a week and start copying whichever group got there first. Grade the prediction and the reasoning, and put the most weight on what changed after it broke.

Tell them what you’re doing, too, early in the term. That was Deslauriers’ own recommendation and it costs nothing. The effort is the mechanism. Most students can hear that and adjust. The ones who work it out alone tend to get it backwards.

The first lab of the year will be ugly. Around the fourth or fifth, the hands stay down longer, and you’ll catch groups arguing with each other about the data before it occurs to them to ask you. Nothing on the state assessment picks that up. Somebody should probably write a rubric for it.

Mentis Sciences builds composite structures and test equipment out of Warner, New Hampshire, mostly for defense and aerospace work. The MSET came out of that side of the shop rather than a curriculum department. The experiment list and the full set of procedures are posted at mentissciences.com.


Bookmark & Share