Popper's question was prompted by a contrast he noticed as a young man in Vienna. Einstein's general relativity made a risky prediction — starlight would bend by a specific amount near the sun — which observation could have flatly contradicted. Other theories he encountered seemed able to explain any behaviour whatsoever after the fact, and their supporters treated this universal applicability as their great strength. Popper concluded it was their fatal weakness.
The insight is that explanatory power and scientific merit can pull in opposite directions. A theory that accommodates every possible outcome has told you nothing about which outcome to expect. Real content consists in what a theory forbids: general relativity would have been in serious trouble if the light had not bent, and that vulnerability is precisely what made the confirmation meaningful when it came.
This reframes scientific method. On the naive picture, scientists gather observations and generalise. Popper denied that induction plays any justificatory role at all: scientists propose bold conjectures, then subject them to severe attempted refutations. A theory that survives serious attempts to kill it is corroborated — not proven, merely not yet refuted. All scientific knowledge stays provisional, which Popper regarded as a strength rather than an embarrassment. It also neatly sidesteps Hume's problem of induction, since no inductive inference is being relied upon.
The criterion has genuine difficulties, and honesty requires stating them. The most serious is the Duhem-Quine problem: hypotheses are never tested in isolation, but always alongside auxiliary assumptions about instruments, background conditions and initial states. When a prediction fails, logic alone does not tell you which element to reject — and rejecting an auxiliary assumption is often the correct scientific move rather than a dodge. The anomalous orbit of Uranus did not refute Newtonian mechanics; it led to the discovery of Neptune. The same manoeuvre that saved Newton could be used to protect a bad theory indefinitely, and falsifiability alone does not distinguish the cases.
There are further problems. Probabilistic and statistical claims are not strictly falsifiable by any finite observation. Some sciences — evolutionary biology, cosmology, historical geology — are heavily explanatory and only indirectly testable, yet clearly scientific. And in practice scientists do not abandon a productive theory at the first anomaly, nor should they. The current view is that falsifiability captures something real and important about scientific responsibility, without being the sharp dividing line Popper hoped for. Demarcation is probably a matter of degree and of research practice rather than a single logical test — but "could anything show this is wrong?" remains among the most useful questions you can ask of any claim.