Building the European Observatory (4/4): Reflections, reckoning and the road ahead
Fenton: Nearly three decades on, what has the Observatory become?
Figueras: Something much larger than we imagined. It began as a fragile experiment with three years of funding and became an institution with real impact on how governments think about health systems. The HiTs evolved into health system reviews and performance assessment tools; studies, policy briefs and rapid responses are widely used; the model inspired observatories elsewhere; and the networks of country experts, policy partners, academics and practitioners became a genuine intellectual community.
Fenton: And what was the hardest thing about holding it together?
Figueras: The partnership model itself. Governments, national institutions, universities and international agencies have different incentives and cultures. But those differences are a strength if you can hold them together. Maintaining that combination through changes of government, funding crises and political pressure was the hardest thing we achieved. I am fully confident the Observatory will not only continue but thrive.
Mossialos: I would add one point about the future. The Observatory should keep its edge. It should remain close enough to governments to be useful, but independent enough to say uncomfortable things. That balance is what made it valuable in the first place.
Fenton: Josep, who picks this up?
Figueras: That is what gives me real confidence. Ewout van Ginneken, who has succeeded me as Director, is outstanding, rigorous, internationally respected and deeply committed to the partnership model. Suszy Lessof as Deputy Director continues to be the institutional backbone. Dimitra Panteli, Matthias Wismar, Jonathan Cylus, Dheepa Rajan and many others now lead major streams of work. With the hubs in Brussels, London and Berlin, the country networks and the wider community of partners and experts, the institutional depth is genuinely there.

***
Fenton: But it was not always easy within your own institutions, was it?
Mossialos: No. We went through rough times even within universities, because this work was not seen as purely academic. Producing accessible evidence for policymakers could be treated as a lesser form of scholarship. Some colleagues questioned whether Observatory publications should count in research assessments. We were often too applied for academics and too academic for bureaucrats.
Fenton: Do academic incentive structures work against this kind of engagement?
Mossialos: Completely. University rankings, promotion criteria and research assessments still revolve largely around journal citations. Policy impact is harder to capture. When you spend months advising a government, very little of that appears in a conventional academic record.
I will be honest: there were times I questioned that trade-off. Every hour advising a government was an hour not spent writing for a leading journal. But then you would see a recommendation reflected in policy, or hear a minister say the evidence had changed how they thought about a problem, and you were reminded why the work mattered.
Things have changed somewhat. The UK Research Excellence Framework now includes impact, and funders increasingly ask what difference research made beyond academia. The Observatory was doing that kind of work before it became fashionable. But the deeper problem has not gone away. Promotion panels still reward papers in high-impact journals, while policy impact remains harder to recognise. Younger colleagues feel this most acutely: they want to do policy-relevant work, but they also need to build academic careers. That is one reason this kind of work remains fragile, and why it still depends on people willing, at certain moments, to trade some conventional academic reward for influence on real decisions.
***
Fenton: One thing that comes through in everything you have described is speed. EU presidencies calling with months to deliver, governments wanting answers to live policy questions. That is not the normal rhythm of academic work. How do you manage the tension between rigour and urgency?
Mossialos: You have to accept that you are not always writing the definitive study. You are providing the best available evidence at the moment a decision has to be made. A presidency lasts six months; a minister may have only a short window to get something through cabinet. If you say "come back in two years when the systematic review is finished," they will decide without you. So the question is not whether the advice is perfect, but whether it is better than what the minister would otherwise have.
There are risks. You can get something wrong or miss a nuance that matters. But the greater risk is silence, leaving policymakers to rely on ideology, anecdote or lobbying. Our job is not to deliver certainty. It is to reduce the uncertainty policymakers face, and to do so on their timeline, not ours.
Figueras: Speed is also a muscle you develop. You can be fast and rigorous if you have already invested in understanding the evidence base. When a country called, we did not start from zero. We had years of accumulated knowledge, expert networks and studies behind us. The rapid response was only the visible part. Underneath it lay years of groundwork.
***
Fenton: This raises a broader methodological question. In medicine, the randomised controlled trial is treated as the gold standard. But in health systems policy, you cannot easily randomise countries. So what counts as evidence in your work?
Mossialos: This is fundamental. Angus Deaton and Nancy Cartwright have argued that RCTs are often given too much authority, that a trial can tell you what happened in a specific setting, but not necessarily why, and that external validity is much harder than people assume. In health systems, this is obvious. You cannot randomise a country into adopting diagnosis-related groups or choosing one financing model over another. These decisions are too large, too political and too embedded in institutional context.
For the biggest questions, evidence usually comes from structured comparative analysis: multiple countries facing similar problems, with careful attention to institutional, political and economic context. We try to identify patterns and mechanisms, under what conditions does a reform work, where has it failed, and why? That is often closer to comparative political science than to clinical trial methodology, and more useful to a health minister deciding among real-world options.
Figueras: And even when trials are feasible in health services, say, testing a new way of organising primary care in a district, the results are almost always context-dependent. What works in a well-funded Swedish municipality may not work in a resource-constrained Romanian region. So we do not just ask, "Did this intervention work?" We ask, "Under what conditions, in what institutional setting, with what resources and political support?" That is the far more useful question for policymakers. It is also why policy dialogue matters. The success of an Observatory policy dialogue depends not only on the quality and relevance of the evidence, but on understanding the policy environment and tailoring the discussion to it.
Mossialos: I would add something here, because it sometimes attracts criticism from colleagues in academia. It is more comfortable to be critical from the outside than to build solutions from inside the room. Critique is essential, it puts pressure on governments, exposes failure, protects the public interest. But critique alone is not enough. Someone has to do the slower, less rewarded work of designing what comes next, sitting with officials who have to implement it, and adjusting it when reality bites. That is where the Observatory chose to operate, and it carries a real risk: the closer you work with governments, the easier it is for your independence to be questioned. Helping to build solutions does not require giving up a critical perspective. It requires holding the two together, which is harder than doing either one alone.
***
Fenton: Let me put a practical question to you. A former student calls and says: I have just been asked to advise a government on health system reform. What should I read first?
Mossialos: I would tell them not to read a health systems book. Not first. Read a good history of the country, its politics, its culture, how power works there. A health system does not exist in a vacuum. The Beveridge model did not emerge because someone designed an optimal financing structure. It emerged from wartime solidarity in Britain, from a particular political moment. The Bismarck model emerged from completely different political calculations. If you walk into a ministry without understanding why the system looks the way it does, you will produce advice that is technically correct and practically useless. The technical literature comes second. You start with context, because context determines whether any reform will fly or crash.
Figueras: I agree, and this connects to how the Observatory differs from consulting firms. Large consultancies often arrive with a standard methodology and confident recommendations based on international best practice. Sometimes that is useful; often it is not enough. Health system reform is not simply a technical problem with a technical answer. It is a political and institutional problem that requires judgement, local understanding and humility. We do not arrive with ready-made answers. We arrive with evidence, options and the experience of what has happened elsewhere. We try to clarify trade-offs. That is less reassuring to a politician who wants certainty, but it is more likely to lead to a reform that will stick.
Mossialos: That is also why the Observatory has lasted. We build long-term relationships. We come back. We are a partner, not a contractor.
***
Fenton: We have heard a lot about the successes. But have there been moments where the work simply did not land, or where you faced outright hostility?
Mossialos: Yes, and it would be dishonest to pretend otherwise. You are entering someone else's territory. Not every politician welcomes that. A defensive government, or a minister who has already decided what to do, creates a very different room. And you often do not know which room it is until you are in it.
I have been invited to present to a minister and before I have even started received the pushback: "Why are you here? There is nothing for us to learn from you." In that moment you could challenge the minister in front of officials, and the meeting would be over. So instead you absorb it. You say: "Of course, Minister, you know your country far better than I do. But I can share what a few other countries have done when facing similar questions, if that is useful." You give them the option to re-engage on their own terms. Sometimes it works. Sometimes it does not. I have had a minister walk out of the room. You learn from it, but you accept that not every engagement will succeed.
Figueras: The Observatory has faced similar moments. There were times when a HiT profile or policy study said something a government did not want to hear, perhaps about inefficiency in hospitals or weaknesses in primary care, and the response was not always gratitude. In those moments, the temptation is to soften the message. Careful wording is sometimes appropriate, because precision matters. But there is a line between being fair and being dishonest, and we never wanted to cross it. Another part of impact is identifying the real policy windows, the honeymoon moments when a new government or minister is open to advice, and proactively offering evidence then. That is when you have the greatest chance of shaping a reform before positions harden.
Fenton: Beyond those confrontational moments, have there been cases where the Observatory's advice or analysis turned out to be wrong, or where a reform you supported did not deliver what was expected?
Figueras: We would be lying if we said no. Health system reform is inherently uncertain. You can assemble the best evidence and still be surprised, because implementation is messy, political circumstances change and economic conditions shift. We have seen reforms that looked sound on paper falter because institutional capacity was weak or the political coalition collapsed. And sometimes the most useful thing we did was to help stop a reform that might have damaged the health system. What matters is being honest about limits. We do not claim to predict the future. We try to set out evidence, options and trade-offs as clearly as we can. When something does not work, we go back and ask why. Failure in one country becomes a lesson for the next.
Mossialos: And sometimes the measure of success is not that you solved the problem but that you stopped it from being ignored. AMR is a case in point: four EU presidencies, nearly two decades of work, specific recommendations implemented, but the core problem is still not solved. The evidence was right. Some of it was taken up. But the structures that make antibiotic development unattractive are deeply entrenched. That is humbling. But without the work, the situation would be worse.
***
Fenton: One final question. AI can now synthesise comparative health systems data in seconds, generate policy briefs and analyse reform options across dozens of countries. Does the Observatory still have a role?
Figueras: More than ever. AI will make some analytical work faster, and the Observatory will use it. But a minister facing a difficult reform does not need a faster literature review. They need to know which reforms have been tried in comparable settings, why they succeeded or failed, and how local politics will shape the options. That requires something AI cannot yet replicate: accumulated institutional knowledge combined with human judgement about what will actually land.
Mossialos: There is also a structural problem. AI systems are trained disproportionately on material from a small number of high-income countries. Ask them to advise on reform in Georgia or Indonesia and they may default to patterns drawn from very different settings. The risk is a confident recommendation grounded in someone else's reality. The Observatory exists because one size does not fit all.
And there is something AI cannot do at all, which is build the relationships that make policy work travel. Ministers and civil servants take advice from people they trust, and trust is built over years of showing up, getting things wrong, and being honest about it. That is the institutional memory the Observatory carries. It is not a database of past reports. It is a network of people who have argued with each other for two decades about what good evidence looks like and what to do with it. That is why the Observatory's role is not diminishing. It is becoming more important.
***
Fenton: Josep, any final thoughts?
Figueras: That institutions are made by people, not by organograms. You can have the best design in the world on paper, but if the people do not believe in it, it will not work. What made the Observatory at the outset was that a small group of people, Elias, Martin, Suszy, Reinhard, Richard and many others believed that evidence could make health systems better, and were willing to devote years to proving it. But it was also the people in the Steering Committee: they worked through bureaucratic and political complexities in their own institutions to keep supporting us. We did not always agree, but the debate and consensus-building made the Observatory stronger. The other lesson is the value of being useful: when something you have written changes how a minister thinks about a problem, that may matter more than a journal citation.
Fenton: Elias, a final reflection before we close?
Mossialos: If I had to add one reflection of my own, it would be this. The hardest thing in this field is not generating evidence. It is keeping intellectual honesty when the political pressure points the other way. Across thirty years I have watched governments reach conclusions before the analysis was done. The discipline of saying what the evidence actually shows, even when it is unwelcome, is harder than it sounds.
But there is a second concern that worries me more, and it concerns my own community. Academic health policy can become too risk-averse. Too much work is incremental. The genuinely big questions, how to finance health systems under demographic pressure, govern health in a fragmenting international order, respond to climate change, use AI responsibly, address structural inequality, require bold, integrative and sometimes unfashionable thinking. The Observatory should remain one of the places where those questions can be asked seriously.
That is what I hope the next generation will carry forward, in the Observatory and beyond: the willingness to ask the questions that actually matter and to follow the evidence even when the field would prefer something tidier. I have not always managed it. That is the discipline I would most want passed on, not as an ideal I embodied, but as one I kept falling short of and returning to.
Figueras: Let me end where Elias began, with the people. The Observatory was never a building or a budget line. It was a community of people who believed that good comparative work could make health systems better, and were willing to put in the years to prove it. That community is larger now than it has ever been, and it extends well beyond the names we have managed to mention here. If the next generation keeps faith with that, the rest will follow.
This is Part 4 of a four-part series on the founding and development of the European Observatory on Health Systems and Policies. Read the other editions here.