
University of Pennsylvania
De La Fuente Lab
AI & Drug Discovery
Dr. César de la Fuente is a Presidential Associate Professor with appointments in the Departments of Bioengineering, Chemical and Biomolecular Engineering, Microbiology, and Psychiatry at the University of Pennsylvania. He earned his PhD in Microbiology and Immunology from the University of British Columbia and a Licenciatura in Biotechnology from the University of León, Spain.
The Extinctome
Most antibiotics we use today were found the same way Alexander Fleming found penicillin: by sampling nature and hoping. At the University of Pennsylvania, Professor César de la Fuente and his Machine Biology Group are using artificial intelligence to search proteins, extinct genomes, and microbial communities for antibiotics evolution has already built, and why waiting for nature to hand over the next candidate was never going to be fast enough.
By Joshua Mullen
In the summer of 1928, a Scottish bacteriologist named Alexander Fleming left his lab for a two-week vacation without bothering to clean up his cluttered bench. He came back to find a stray blue-green mold growing in one of his forgotten petri dishes, and circling that mold, a ring of dead bacteria, colonies that had been thriving everywhere else on the plate, wiped out wherever the mold's influence reached. Fleming had stumbled onto penicillin, and with it, the founding method of antibiotic discovery: expose bacteria to something from the natural world, and see what survives.
That method worked remarkably well for decades, but stumbling has its limits, and the easy accidents ran out long ago. Antibiotic-resistant infections are now on pace to kill an estimated 10 million people a year within a generation, roughly one death every three seconds, while the traditional search, screening soil bacteria and hoping for another lucky mold, has slowed to a crawl. At the University of Pennsylvania, Professor César de la Fuente and his Machine Biology Group are trying to replace luck with search, using artificial intelligence to comb through proteins, extinct genomes, and unexplored microbial communities for antibiotics evolution has already built, without needing an accident to find them.
A Searchable Space, Not a Known List
Even inside the world of AI-assisted discovery, most approaches have quietly kept one of Fleming's old assumptions: that a useful molecule comes from somewhere specific, an immune-related protein, a known antimicrobial family, an organism already suspected of producing something useful. De la Fuente's approach strips that assumption out almost entirely.
"We can avoid presupposing that an antibiotic must originate from a particular organism, protein family, or immune pathway," he explained. "Instead, the model learns relationships between amino acid sequence and function, including patterns of charge, hydrophobicity, structure, and other physicochemical properties that may enable a peptide to interact with bacterial membranes or intracellular targets." The model isn't asking whether a protein is already labeled antimicrobial. It's asking whether a specific fragment of it has the physical properties needed to punch a hole in a bacterial membrane, regardless of what the rest of the protein actually does.
"A useful model must generalize beyond its training examples," de la Fuente said. "It should recognize functional principles wherever they occur, including within proteins whose known biological roles have nothing to do with immunity." That's exactly what his lab's first major result demonstrated: a 2022 study mining the human proteome found dozens of encrypted antimicrobial peptides hiding inside proteins with no known connection to the immune system at all, uncovering what the lab describes as a previously unrecognized branch of host immunity. He described the underlying principle in a single sentence: AI, in his account, "compresses biological complexity into representations that can reveal similarities in function even when the sequences and their evolutionary origins appear very different."
An Unusually Wide Footprint
De la Fuente's own account of why he approaches the problem this way traces back to how he was trained. He earned his PhD in 2014 from the University of British Columbia and completed postdoctoral research at MIT before joining Penn, where he became one of the youngest tenured professors in the history of Penn Medicine. His faculty appointments alone tell part of the story: he holds primary positions in both the Department of Bioengineering and the Department of Chemical and Biomolecular Engineering in Penn's School of Engineering, plus appointments in Psychiatry and Microbiology in the Perelman School of Medicine, an unusually wide footprint for a single lab to occupy at once.
He's described the philosophy behind that breadth directly. "The boundary among disciplines is fading," he said, describing his lab's approach as built on what he calls "transdisciplinarity," the convergence of chemistry and physics with synthetic biology, microbiology, and computer science. He draws a specific distinction between how biologists and engineers tend to approach the same problem. "Biologists tend to have a holistic view of the world and aim to understand biological systems, whereas engineers see the world from a very practical perspective and aim to build new and useful tools," he said. "I reasoned that they were the perfect raw material for creating useful tools for deciphering diseases." That combination is arguably the real design principle behind everything else in this piece, not a specific algorithm, but a lab built deliberately to sit across a boundary most research groups stay on one side of.
Digging Up Molecules That No Longer Exist
If antimicrobial fragments hide inside living proteins with no obvious connection to immunity, the same logic raises a stranger question: could those hidden molecules also exist in the proteins of organisms that no longer exist at all? That question became the basis for what de la Fuente's lab calls molecular de-extinction, mining a digital archive of extinct-organism sequence data his team calls the "extinctome" for candidate antibiotics.
It's tempting to read this as a flashy idea dreamed up because extinction sounds compelling, but de la Fuente says that isn't how it started. "The idea emerged from a scientific hypothesis rather than from the appeal of extinction itself," he said. "We had found that proteins can contain encrypted peptide sequences with antimicrobial potential, even when the full-length proteins have no recognized immune function. If those hidden molecules are distributed across living proteomes, we reasoned that they should also exist in the proteomes of extinct organisms." Extinct species, in his framing, are evolutionary experiments no longer accessible in the living world, their reconstructed proteomes preserving molecular solutions shaped by environments and pathogens that no longer exist either.
To test the idea at scale, the lab built a deep learning model called APEX, for Antibiotic Peptide de-Extinction, combining recurrent neural networks with attention networks to hunt for these encrypted fragments. Turned loose on more than ten million peptides drawn from hundreds of extinct proteomes, including woolly mammoths, straight-tusked elephants, ancient sea cows, and extinct giant elk, APEX flagged over 37,000 sequences with predicted antimicrobial activity, more than 11,000 of which don't exist in any living organism today. Some of the resulting candidates, among them neanderthalin and mammuthusin, have already shown effectiveness against infection in preclinical mouse models. As de la Fuente put it, "the deeper principle is that evolutionary history constitutes an enormous, largely unexplored archive of molecular diversity."

Antibiotics From the Edge of Life
The same search strategy that worked on human and extinct proteins eventually pointed de la Fuente's lab toward one of life's stranger corners: archaea, single-celled organisms that often thrive in conditions that would kill nearly anything else, extreme heat, crushing pressure, high acidity, high salinity.
"Archaea are not simply unusual bacteria," de la Fuente said. "Their membranes, molecular machinery, and evolutionary history differ fundamentally from those of bacteria and eukaryotes." He's careful about what the resulting data actually shows. "We cannot assume every peptide we identified is naturally produced by archaea specifically as an antibiotic," he said. "Rather, our models uncovered fragments of archaeal proteins that display antimicrobial activity when synthesized and tested," a real distinction between a molecule that works once isolated in a lab and one an organism is actually deploying as a weapon in the wild.
What the search did reveal was genuinely useful: many of the resulting molecules, which the lab named archaeasins, carried less positive electrical charge than conventional antimicrobial peptides. That matters because most known antimicrobial peptides work the way a magnet does, their positive charge pulling them toward the negatively charged surface of a bacterial cell. If that were the only viable mechanism, it would sharply limit how much design space is actually worth searching. "This suggests that strong positive charge is not the only possible solution," de la Fuente said. "Archaeal sequence space may contain alternative combinations of charge, hydrophobicity, and structure that enable antimicrobial activity through different interactions or mechanisms," some of which have already proven effective against drug-resistant bacteria in animal models.
Mining the Microbial Dark Matter
Human proteins, extinct genomes, and archaea aren't the only unlikely places de la Fuente's lab has searched. In a separate large-scale effort, his team computationally analyzed 63,410 metagenomes and 87,920 microbial genomes, sequence data pulled from environments like soil, ocean water, and the human gut, much of it never assigned to any specific known organism at all, a category researchers sometimes call microbial dark matter. That search alone turned up close to a million new candidate antimicrobial molecules, which the lab released publicly rather than holding as proprietary, aiming to let other labs synthesize and test them faster. A related project mining thousands of human gut microbiomes specifically surfaced prevotellin-2, an antimicrobial peptide hiding inside a protein from Prevotella copri, a common gut bacterium with no previously known role in producing antibiotics of its own.
The same underlying logic runs through every branch of this work. Useful molecules aren't concentrated in a handful of well-studied organisms. They're scattered thinly across nearly the entire tree of life, in genomes nobody has had any specific reason to search, and finding them means treating "nobody has looked here yet" as a reason to look, rather than a reason to skip it.
What a Model Can't Tell You
None of this search matters if a computational prediction and a working antibiotic get treated as the same thing, and de la Fuente is unusually direct about where that line sits.
Deciding which predictions are worth testing comes down to weighing several factors together: "the model's confidence, the molecule's novelty, predicted potency and selectivity, physicochemical properties, synthetic feasibility, and whether it can answer an interesting biological or therapeutic question." His lab doesn't only chase the model's most confident predictions, either. "We also deliberately test some candidates that challenge the model's assumptions," he said, "because informative failures can be as valuable for improving a model as successful predictions," treating a wrong prediction as data rather than wasted effort.
The clearest statement of his philosophy might be this one: "A prediction tells us that a molecule is worth testing; it does not establish that the molecule is an antibiotic." Everything after that still has to happen with real bacteria and real cells. "Laboratory experiments reveal whether it actually inhibits or kills bacteria, which pathogens it affects, whether it harms human cells, how stable it is, how it works, and whether it remains effective in an animal model." The model's job is to shrink an impossibly large search space down to a short list. The lab's job is everything after that.

Faster Discovery Isn't Enough
It would be easy to tell this as a simple story of computers finding cures faster than humans ever could. De la Fuente's own account is more measured.
"AI has changed what appears technically possible," he said. "Traditional antibiotic discovery often depends on slow, local sampling and large-scale brute-force screening. Computational methods allow us to search millions, or even billions, of sequences and identify candidates for experimental testing in hours." Across his lab's projects, that shift has compressed work that once took years into a process completed in hours, an estimated speedup of several million-fold. But he doesn't let the speed gain stand in for a solved problem. "Faster discovery alone will not solve the antibiotic crisis," he said. "AI accelerates the front end of the process, but promising molecules must still undergo optimization, toxicology, manufacturing, clinical testing, and regulatory review." He points to something structural, too: antibiotics face an economic bind most other drugs don't. "Society needs new antibiotics to be available, but responsible stewardship requires that they be used sparingly." A company can't profit much from a medicine doctors are actively trying not to prescribe, which is exactly what good stewardship demands, and it's a large part of why so few companies invest in the space regardless of how good the science gets. "The field therefore needs sustained investment, stronger translational infrastructure, and incentives that make antibiotic development economically viable," he said. "AI can help refill the pipeline, but the surrounding scientific, clinical, and economic system must be capable of carrying those discoveries forward."
From Finding to Designing
Everything described so far is a search problem: sift through biological data that already exists and find the antimicrobial molecules hiding inside it. De la Fuente's lab has increasingly moved toward a more ambitious task.
"Discovery asks, 'What has evolution already made?'" he said. "Design asks, 'What molecule should we make to achieve a particular objective?'" A discovery model is a detector, scanning existing sequences for a signal that's already there. A design model is a generator, proposing sequences that may never have existed anywhere in nature, built to hit a target specified in advance, rather than waiting for evolution to have solved the same problem somewhere across three billion years of biological history. "Natural sequences provide valuable evolutionary starting points," de la Fuente explained, "whereas generative models can explore combinations that may never have existed. We can specify multiple desired properties, such as potency against a particular pathogen, low toxicity toward human cells, stability, or ease of synthesis, and computationally optimize them together." The catch: "these objectives can conflict. Increasing antimicrobial potency, for example, can also increase toxicity."
His lab's next step in this direction is a model called ApexOracle, designed to analyze a new pathogen, identify its genetic weaknesses, match it against candidate antimicrobial peptides, and predict how a resulting antibiotic would actually perform in lab testing, a system de la Fuente has described as one that "converges understanding in chemistry, genomics, and language." It's a fitting place for the work to be heading. De la Fuente has spent his career treating biology, living, extinct, and microbial, as a vast, searchable archive of molecular solutions. The next chapter isn't just about searching that archive more thoroughly. It's about learning to write new entries.