Roger Mark (left) and George Moody are two of the original founding members behind PhysioNet, which was based on the MIT–BIH Arrhythmia Database developed in the late 1970s. Last year, more than 15,000 scientific publications cited PhysioNet, and users from more than 180 countries have registered on the platform.
Credits:
Photo courtesy of the Harvard-MIT Program in Health Sciences and Technology.
The visionary PhysioNet platform launched 25 years ago, based on a system developed at MIT in the 1970s. It has become one of the most comprehensive biomedical and clinical data repositories in existence.
Emma Foehringer Merchant | Institute for Medical Engineering and Science
Before the advancement of scientific data storage and collaboration via the cloud, medical investigators seeking health research breakthroughs had to overcome significant obstacles to collaboration and key clinical data gathering.
Data were siloed and difficult to distribute, so those looking to undertake research had no option but to gather them themselves. This not only made research more expensive, but it was challenging to compare findings across datasets.
In 1975, researchers studying arrhythmias at MIT and Boston’s Beth Israel Hospital envisioned another way: the team began collecting and digitizing electrocardiogram recordings with the intention of not only studying them, but of also making them available to the wider research community.
The team built their own computers for the process, painstakingly duplicated tapes one by one, and created more than 100,000 annotations for the recordings. The process took years, but by summer 1980, the tapes were finally ready. The team initially thought their tool would reach fewer than a dozen academic and industry groups. But interest kept pouring in. Over the next decade, they went on to mail about 100 copies.
The data eventually became the first database of the global platform PhysioNet — founded in 1999 at the Harvard-MIT program in Health Sciences and Technology (HST) — as a clinical data repository for complex physiological signals.
At the time, that type of data-sharing, which may seem like the default today, was a near-revolutionary idea. PhysioNet’s “founding was incredibly visionary,” says Thomas Heldt, Richard J. Cohen (1976) Professor in Medicine and Biomedical Physics, associate director of MIT’s Institute for Medical Engineering and Science (IMES), and the senior author of a recent paper in Nature Health examining the platform’s impact.
Eventually, those magnetic tapes sent through the mail became burned CD-ROMs, which then evolved into FTP servers hosted on the newly minted internet. Today, as PhysioNet looks back at over 25 years of operation, the platform hosts hundreds of databases, and has become one of the most comprehensive biomedical and clinical data repositories in existence. Last year, more than 15,000 scientific publications cited PhysioNet, and users from more than 180 countries have registered on the platform. It is widely used by researchers, manufacturers, and clinical decision-makers.
“The research impact is truly significant,” says Heldt, who is also a professor in the MIT Department of Electrical Engineering and Computer Science and a principal investigator at the Research Laboratory of Electronics, “and quite humbling.”
“It is really beautiful to see that such a vision has proven right and so enabling for so many people.”
Setting a standard
Around 2009, a PhD student named Tom Pollard was conducting research on critically ill patients at one of London’s leading hospital systems. Although the hospital generated large volumes of valuable clinical data, the infrastructure and processes needed to curate and support their wider research use were still developing.
“Hospital data were collected primarily to support immediate patient care, with less attention given to how they might be curated and reused for research,” says Pollard, now a research scientist at MIT’s Laboratory for Computational Physiology (LCP), technical director of PhysioNet, and the lead author on the Nature Health paper.
The problem was not simply privacy. Hospital information systems were built primarily to support patient care and administration, not research. Data were fragmented across systems and rarely curated with future reuse in mind, making it difficult and expensive to turn them into coherent research resources.
But Pollard needed data to complete his dissertation. After poking around on the internet, he eventually discovered the Medical Information Mart for Intensive Care (MIMIC), a database of de-identified electronic health records hosted by PhysioNet. Recognizing its potential, his clinical supervisor, Kevin Fong, organized a visit to Boston. Soon afterward, Fong and Pollard were sitting across the table from Roger Mark, discussing how their teams might collaborate.
Academic incentives have long favored publications and exclusive analyses over the less-visible work involved in preparing data for others to use. That tension persists today. PhysioNet’s founders embraced a different model, believing that sharing research resources could accelerate discovery and ultimately improve human health, he says. MIMIC became central to Pollard’s dissertation, and after completing his PhD, he came to MIT to help build the next generation of the database.
In the years since PhysioNet was established, the value of sharing research data has gained much wider recognition. The late Roger Mark, MIT’s distinguished professor of health sciences and technology emeritus and one of PhysioNet’s founders, described its purpose as building an “accessible multinational community around data” to “positively impact global health.”
Earlier this year, Mark and the late George Moody, PhysioNet’s co-founder, jointly received the prestigious IEEE Biomedical Engineering Award for their contributions to PhysioNet and biomedical signal processing. IEEE cited their “leadership in ECG signal processing and global dissemination of curated biomedical and clinical databases, thereby accelerating biomedical research worldwide.”
The source code for the platform, like much of its data, is public. According to the Nature piece: “As the platform evolved, PhysioNet’s community broadened substantially beyond its origins in signal processing and cardiovascular health to encompass clinical informatics, critical care and machine learning for health.” People have used that to build their own PhysioNet-esque infrastructure, says Heldt. Pollard points to similar platforms like Health Data Nexus as examples of PhysioNet’s legacy.
Although there are now more resources out there hosting similar electronic health data, according to Google DeepMind researcher Vivek Natarajan, both PhysioNet and MIMIC “set the standard,” he says, “and it’s still the standard right now.”
That standard, according to those who use the platform, changed how research is conducted. Access to data should not be the determinant for which ideas are possible, according to Ziad Obermeyer, an associate professor at the University of California at Berkeley School of Public Health and the College of Computing, Data Science, and Society.
“PhysioNet changed how I think about the bottleneck in research. It is often not ideas or talent. It is friction. When access to data is slow, expensive, and hard, the ideas that die first are the high-risk ones, the things that probably will not work, but would be transformative if they did. That is exactly the wrong model if you want real progress,” he says. “PhysioNet lowers the fixed cost of trying ambitious ideas, and that changes what science becomes possible.”
The AI boom
PhysioNet, once a repository mainly for those working in biomedical signal processing and the health-care fields, has evolved in its 25 years. Originally, the holdings consisted solely of cardiovascular ECG data. Now PhysioNet is a largely a source for electronic health records, imaging data, and software and AI models.
Particularly as artificial intelligence approaches took off, “the community shifted,” Heldt explains. Those in need of signal processing data still use PhysioNet databases, but the pool of users has expanded to encompass staff at large tech companies, teachers, and practitioners in all areas of medicine, as well as researchers in health-related machine learning and AI. Today, that latter group “dominates the user community,” says Heldt.
The platform hosts the highest-quality datasets available for health-care AI research, according to Natarajan, whose research involves AI, science, and medicine and who has published several papers that used its datasets.
“It has been an important cornerstone that has catalyzed all the progress in health-care AI over the last decade,” says Natarajan. In addition to using PhysioNet data, he and his colleagues have contributed data to the platform, helping create the self-sustaining ecosystem that typifies PhysioNet.
Looking toward the coming decades, stewards of the platform like Heldt and Pollard envision continuing to expand its reach with an annual conference. The team is also preparing to pilot a new system that will allow users to annotate data and contribute their own expertise, enriching PhysioNet’s resources for the next phase of the platform.
“The kind of research that people want to do now needs to be interdisciplinary. Statisticians, computer scientists, clinicians, pharmacists, and nurses must all come together and contribute their knowledge to develop algorithms that are useful for people” says Pollard. “The community has broadened, and advances in AI have expanded both the questions researchers can address and what they believe is possible.”
*Originally published in MIT News.