rxrx-20241231
Securities registered pursuant to Section 12(b) of the Act:
Title of each class Trading symbol(s) Name of each exchange on which registered
Class A Common Stock, par value $0.00001 RXRX Nasdaq Global Select Market
Large accelerated filer x Non-accelerated filer ☐
Accelerated filer ☐ Smaller reporting company ☐
Emerging growth company ☐
>$450 MillionIN COLLABORATION PAYMENTS TO DATE
Table of Contents
TABLE OF CONTENTS
PART 01 Item 1. Business 10
Item 1A. Risk Factors 99
Item 1B. Unresolved Staff Comments 161
Item 1C. Cybersecurity 161
Item 2. Properties 162
Item 3. Legal Proceedings 163
Item 4. Mine Safety Disclosures 163
Item 6. [Reserved] 165
Item 7A. Quantitative and Qualitative Disclosures About Market Risk 177
Item 8. Financial Statements and Supplementary Data 179
Item 9A. Controls and Procedures 220
Item 9B. Other Information 222
Item 9C. Disclosure Regarding Foreign Jurisdictions that Prevent Inspections 222
PART 03 Item 10. Directors, Executive Officers and Corporate Governance 223
Item 11. Executive Compensation 223
Item 14. Principal Accounting Fees and Services 223
PART 04 Item 15. Exhibits and Financial Statement Schedules 224
3
Table of Contents
PART I
Risk Factor Summary
Below is a summary of the principal factors that make an investment in the common stock of Recursion Pharmaceuticals, Inc. (Recursion, the Company, we, us, or our) risky or speculative. This summary does not address all of the risks we face. Additional discussion of the risks summarized below, and other risks that we face, can be found in the section titled “Item 1A. Risk Factors” in this Annual Report on Form 10-K.
Risks Related to Our Limited Operating History, Financial Position, and Need for Additional Capital
◦We are a clinical-stage biotechnology company with a limited operating history and no products approved by regulators for commercial sale, which may make it difficult to evaluate our current and future business prospects
◦We have incurred significant operating losses since our inception and anticipate that we will incur continued losses for the foreseeable future.
◦We will need to raise substantial additional funding. If we are unable to raise capital when needed, we would be forced to delay, reduce, or eliminate at least some of our product development programs, business development plans, strategic investments, and potential commercialization efforts, and to possibly cease operations.
◦Raising additional capital and issuing additional securities may cause dilution to our stockholders, restrict our operations, require us to relinquish rights to our technologies or drug candidates, and divert management’s attention from our core business.
◦We may be required to repurchase for cash all, or to facilitate the purchase by a third party of all, of the shares of Class A common stock that were issued to the Bill & Melinda Gates Foundation, or the Gates Foundation, in exchange for the Exscientia American Depositary Shares (ADSs) that the Gates Foundation purchased from Exscientia in an October 2021 private placement if we default under the global access commitments agreement with Exscientia, which could have an adverse impact on us.
◦We are engaged in strategic collaborations and we intend to seek to establish additional collaborations, including for the clinical development or commercialization of our drug candidates. If we are unable to establish collaborations on commercially reasonable terms or at all, or if current and future collaborations are not successful, we may have to alter our development and commercialization plans.
◦We have no products approved for commercial sale and have not generated any revenue from product sales. We or our current and future collaborators may never successfully develop and commercialize our drug candidates, which would negatively affect our results of operation and our ability to continue our business operations.
◦If we engage in future acquisitions or strategic partnerships, this may increase our capital requirements, dilute our stockholders’ equity, cause us to incur debt or assume contingent liabilities, and subject us to other risks.
Risks Related to the Discovery and Development of Drug Candidates
◦Our approach to drug discovery is unique and may not lead to successful drug products for various reasons, including, but not limited to, challenges identifying mechanisms of action for our candidates.
◦Our drug candidates are in preclinical or clinical development, which are lengthy and expensive processes with uncertain outcomes and the potential for substantial delays.
◦If we experience delays or difficulties in the enrollment of patients in clinical trials, our receipt of necessary regulatory approvals could be delayed or prevented.
◦Our planned clinical trials, or those of our current and potential future collaborators, may not be successful or may reveal significant adverse events not seen in our preclinical or nonclinical studies, which may result in a safety profile that could inhibit regulatory approval or market acceptance of any of our drug candidates.
◦We may develop drug candidates for use in combination with other therapies, which exposes us to additional risks.
◦We conduct clinical trials for our drug candidates outside the United States, and the FDA and similar foreign regulatory authorities may not accept data from such trials.
◦It is difficult to establish with precision the incidence and prevalence for target patient populations of our drug candidates. If the market opportunities for our drug candidates are smaller than we estimate, or if any approval that we obtain is based on a narrower definition of the patient population, our revenue and ability to achieve profitability will be adversely affected, possibly materially.
◦We may never realize a return on our investment of resources and cash in our drug discovery collaborations.
◦We face substantial competition, which may result in others discovering, developing, or commercializing products before, or more successfully than, we do.
◦Because we have multiple programs and drug candidates in our development pipeline and are pursuing a variety of target indications and treatment modalities, we may expend our limited resources to pursue a particular drug candidate and fail to
4
Table of Contents
capitalize on development opportunities or drug candidates that may be more profitable or for which there is a greater likelihood of success.
◦Our product candidates may cause significant adverse events, toxicities or other undesirable side effects when used alone or in combination with other approved products or investigational new drugs that may result in a safety profile that could prevent regulatory approval, prevent market acceptance, limit their commercial potential or result in significant negative consequences.
Risks Related to Our Platform and Data
◦We have invested, and expect to continue to invest, in research and development efforts to further enhance our drug discovery platform, which is central to our mission. If the return on these investments is lower or develops more slowly than we expect, our business and operating results may suffer.
◦Our information technology systems and infrastructure may fail or experience security breaches and incidents that could adversely impact our business and operations and subject us to liability.
◦Interruptions in the availability of server systems or communications with internet or cloud-based services, or failure to maintain the security, confidentiality, accessibility, or integrity of data stored on such systems, could harm our business.
◦Our solutions utilize third-party open source software (OSS), which presents risks that could adversely affect our business and subject us to possible litigation.
◦Issues relating to the use of artificial intelligence and machine learning in our offerings could adversely affect our business and operating results.
Risks Related to Our Operations/Commercialization
◦Even if any drug candidates we develop receive marketing approval, they may fail to achieve the degree of market acceptance by physicians, patients, healthcare payors, and others in the medical community necessary for commercial success.
◦If we are unable to establish sales and marketing capabilities or enter into agreements with third parties to sell and market any drug candidates we may develop, we may not be successful in commercializing those drug candidates, if and when they are approved.
◦We are subject to regulatory and operational risks associated with the physical and digital infrastructure at both our internal facilities and those of our external service providers.
◦The manufacture of drugs is complex, and our third-party manufacturers may encounter difficulties in production or supply chain. If any of our third-party manufacturers encounter such difficulties, our ability to provide adequate supply of our product candidates for clinical trials or our products for patients, if approved, could be delayed or prevented.
Risks Related to Our Intellectual Property
◦Our success significantly depends on our ability to obtain and maintain patents of adequate scope covering our proprietary technology and drug candidate products. Obtaining and maintaining patent assets is inherently challenging, and our pending and future patent applications may not issue with the scope we need, if at all.
◦Our current proprietary position for certain drug product candidates depends upon our owned or in-licensed patent filings covering components of such drug product candidates, manufacturing-related methods, formulations, and/or methods of use, which may not adequately prevent a competitor or other third party from using the same drug candidate for the same or a different use.
◦We may not be able to protect our intellectual property and proprietary rights throughout the world.
◦If we do not obtain patent term extension and data exclusivity for any drug product candidates we may develop, our business may be materially harmed.
◦We may need to license certain intellectual property from third parties, and such licenses may not be available or may not be available on commercially reasonable terms.
◦Changes in U.S. patent law could diminish the value of patents in general, thereby impairing our ability to protect our products.
◦Issued patents covering our drug product candidates and proprietary technology that we have developed or may develop in the future could be found invalid or unenforceable if challenged in court or before administrative bodies in the United States or abroad.
Risks Related to Acquisitions
◦The failure to integrate successfully the businesses of Recursion and Exscientia in the expected timeframe would adversely affect Recursion’s future business and financial performance.
◦The anticipated benefits of the business combination with Exscientia may vary from expectations.
◦As a company with substantial operations outside of the United States, we are subject to economic, political, regulatory and other risks associated with international operations.
5
Table of Contents
Risks Related to Government Regulation
◦We may be unable to obtain U.S. or foreign regulatory approval and, as a result, may be unable to commercialize our product candidates.
◦The FDA, EMA and other comparable foreign regulatory authorities may not accept data from trials conducted in locations outside of their jurisdiction.
◦Even if we receive FDA or other regulatory approval for any of our drug candidates, we will be subject to ongoing regulatory obligations and other conditions that may result in significant additional expense, as well as the potential recall or market withdrawal of an approved product if unanticipated safety issues are discovered.
◦The FDA and other regulatory agencies actively enforce the laws and regulations prohibiting the promotion of off-label uses.
◦Though we have been granted orphan drug designation for certain of our drug candidates, we may be unsuccessful or unable to maintain the benefits associated with such a designation, including the potential for market exclusivity.
◦We are subject to U.S. and foreign laws regarding privacy, data protection, and data security that could entail substantial compliance costs, while the failure to comply could subject us to significant liability.
◦Regulatory and legislative developments related to the use of AI could adversely affect our use of such technologies in our products, services, and business.
Other Risks
◦Third parties that perform some of our research and preclinical testing or conduct our clinical trials may not perform satisfactorily or their agreements may be terminated.
◦Third parties that manufacture our drug candidates for preclinical development, clinical testing, and future commercialization may not provide sufficient quantities of our drug candidates or products at an acceptable cost, which could delay, impair, or prevent our development or commercialization efforts.
◦We may not realize all of the anticipated outcomes and benefits of our Acquisitions.
◦Our future success depends on our ability to retain key executives and experienced scientists, and to attract, retain, and motivate qualified personnel.
◦We have identified a material weakness in our internal control over financial reporting.
6
Cautionary Note Regarding Forward-Looking Statements
This Annual Report on Form 10-K contains “forward-looking statements” about us and our industry within the meaning of Section 27A of the Securities Act of 1933, as amended, and Section 21E of the Securities Exchange Act of 1934, as amended. All statements other than statements of historical facts are forward-looking statements. In some cases, you can identify forward-looking statements by terms such as “may,” “will,” “should,” “would,” “expect,” “plan,” “anticipate,” “could,” “intend,” “target,” “project,” “contemplate,” “believe,” “estimate,” “predict,” “potential,” or “continue” or the negative of these terms or other similar expressions. Forward-looking statements contained in this report may include without limitation those regarding:
•our research and development programs;
•the initiation, timing, progress, results, and cost of our current and future preclinical and clinical studies, including statements regarding the design of, and the timing of initiation and completion of, studies and related preparatory work, as well as the period during which the results of the studies will become available and key milestones will be met;
•our ability to use our combined assets from our recent business combination to create a fully integrated, technology-first drug discovery platform;
•the timing and likelihood of our ability to shift our wet-lab from a source of data generation to a model for validating data from AI-generated results;
•the ability and willingness of our collaborators to continue research and development activities relating to our development candidates and investigational medicines;
•future agreements with third parties in connection with the commercialization of our investigational medicines and any other approved product;
•the timing, scope, and likelihood of regulatory filings and approvals, including the timing of Investigational New Drug applications and final approval by the U.S. Food and Drug Administration, or FDA, of our current drug candidates and any other future drug candidates, as well as our ability to maintain any such approvals;
•the timing, scope, or likelihood of foreign regulatory filings and approvals, including our ability to maintain any such approvals;
•the size of the potential market opportunity for TechBio companies, including the expected impact of AI-enabled technologies;
•the size of the potential market opportunity for our drug candidates, including our estimates of the number of patients who suffer from the diseases we are targeting;
•our ability to identify viable new drug candidates for clinical development and the rate at which we expect to identify such candidates, whether through an inferential approach or otherwise;
•our expectation that the assets that will drive the most value for us are those that we will identify in the future using our datasets and tools;
•our ability to develop and advance our current drug candidates and programs into, and successfully complete, clinical studies;
•our ability to reduce the time or cost or increase the likelihood of success of our research and development relative to the traditional drug discovery paradigm, including the use of data sets from our partners to accelerate the development of our AI-enabled technologies;
•our ability to improve, and the rate of improvement in, our infrastructure, datasets, biology, technology tools, and drug discovery platform, and our ability to realize benefits from such improvements;
•our ability to effectively use machine learning and artificial intelligence in our drug development process;
•our ability to leverage our collaborations and partnerships to develop our products and grow our business;
•our expectations related to the performance and benefits of our BioHive-2 supercomputer, Recursion OS, and our digital chemistry platform;
•our ability to realize a return on our investment of resources and cash in our drug discovery collaborations;
•our ability to sell or license assets and re-invest proceeds into funding our long-term strategy;
•our ability to scale like a technology company and to add more programs to our pipeline each year;
•our ability to acquire and generate datasets to train and develop our AI-enabled technologies;
•our ability to successfully compete in a highly competitive market;
•our manufacturing, commercialization, and marketing capabilities and strategies;
•our plans relating to commercializing our drug candidates, if approved, including the geographic areas of focus and sales strategy;
•our expectations regarding the approval and use of our drug candidates in combination with other drugs;
•the rate and degree of market acceptance and clinical utility of our current drug candidates, if approved, and other drug candidates we may develop;
•our competitive position and the success of competing approaches that are or may become available, including with respect to our AI-enabled technologies;
•our estimates of the number of patients that we will enroll in our clinical trials and the timing of their enrollment;
7
•the beneficial characteristics, safety, efficacy, and therapeutic effects of our drug candidates;
•our plans for further development of our drug candidates, including additional indications we may pursue;
•our ability to adequately protect and enforce our intellectual property and proprietary technology, including the scope of protection we are able to establish and maintain for intellectual property rights covering our current drug candidates and other drug candidates we may develop, receipt of patent protection, the extensions of existing patent terms where available, the validity of intellectual property rights held by third parties, the protection of our trade secrets, and our ability not to infringe, misappropriate or otherwise violate any third-party intellectual property rights;
•the impact of any intellectual property disputes and our ability to defend against claims of infringement, misappropriation, or other violations of intellectual property rights;
•our ability to keep pace with new technological developments, including with respect to AI;
•our ability to utilize third-party open source software and cloud-based infrastructure, on which we are dependent;
•the adequacy of our insurance policies and the scope of their coverage;
•the potential impact of a pandemic, epidemic, or outbreak of an infectious disease, such as COVID-19, or natural disaster, global political instability, or warfare, and the effect of such outbreak or natural disaster, global political instability, or warfare on our business and financial results;
•our ability to achieve net-zero greenhouse gas emissions across our operations;
•our ability to maintain our technical operations infrastructure to avoid errors, delays, or cybersecurity breaches;
•our continued reliance on third parties to conduct additional clinical trials of our drug candidates, and for the manufacture of our drug candidates for preclinical studies and clinical trials;
•our ability to obtain, and negotiate favorable terms of, any collaboration, licensing or other arrangements that may be necessary or desirable to research, develop, manufacture, or commercialize our platform and drug candidates;
•the pricing and reimbursement of our current drug candidates and other drug candidates we may develop, if approved;
•our estimates regarding expenses, future revenue, capital requirements, and need for additional financing;
•our financial performance;
•the period over which we estimate our existing cash and cash equivalents will be sufficient to fund our future operating expenses and capital expenditure requirements;
•our ability to raise substantial additional funding;
•the impact of current and future laws and regulations, and our ability to comply with all regulations that we are, or may become, subject to;
•the need to hire additional personnel and our ability to attract and retain such personnel;
•the impact of any current or future litigation, which may arise during the ordinary course of business and be costly to defend;
•our ability to maintain effective internal control over financial reporting and disclosure controls and procedures, including our ability to remediate the material weakness in internal control over financial reporting;
•our anticipated use of our existing resources and the net proceeds from our public offerings; and
•other risks and uncertainties, including those listed in the section titled “Risk Factors.”
We have based these forward-looking statements largely on our current expectations and projections about our business, the industry in which we operate, and financial trends that we believe may affect our business, financial condition, results of operations, and prospects. These forward-looking statements are not guarantees of future performance or development. These statements speak only as of the date of this report and are subject to a number of risks, uncertainties and assumptions described in the section titled “Risk Factors” and elsewhere in this report. Because forward-looking statements are inherently subject to risks and uncertainties, some of which cannot be predicted or quantified, you should not rely on these forward-looking statements as predictions of future events. The events and circumstances reflected in our forward-looking statements may not be achieved or occur and actual results could differ materially from those projected in the forward-looking statements. Except as required by applicable law, we undertake no obligation to update or revise any forward-looking statements contained herein, whether as a result of any new information, future events, or otherwise.
In addition, statements that “we believe” and similar statements reflect our beliefs and opinions on the relevant subject. These statements are based upon information available to us as of the date of this report. While we believe such information forms a reasonable basis for such statements, the information may be limited or incomplete, and our statements should not be read to indicate that we have conducted an exhaustive inquiry into, or review of, all potentially available relevant information. These statements are inherently uncertain and you are cautioned not to unduly rely upon them.
8
Table of Contents
Item 1. Business.
Recursion At-a-Glance
The Problem
Discovering and developing effective new medicines is among the most challenging human pursuits due to the incredible complexity of biology, the vastness of chemical space, and both the ineffectiveness and inefficiencies in clinical development. Today more than 90% of clinical trials fail.
Our Mission and Philosophy
Recursion was founded in 2013 to decode biology to radically improve lives by building the next generation biopharma company from the ground-up, leveraging the latest technology tools to help map and navigate the incredible complexity of biology, vastness of chemical space, and the broken drug development process. We believe that rapid commoditization of artificial intelligence creates a once-in-a-generation market opportunity for companies with the ability to build the right datasets in biology and chemistry to win.
Our Competitive Advantage
That’s why we have invested heavily in creating one of the most sophisticated automated wet-laboratories in the world where robots and sensors help us conduct and digitize millions of real-life experiments each week, spanning cellular systems, chemical systems, tissue systems, and animal models. We also partner with select companies to aggregate and relate data from patients and health systems at scale. In addition, we command industry-leading computational capabilities including BioHive-2 (the fastest supercomputer wholly owned and operated by any biopharma company), and employ hundreds of data scientists, software engineers, and AI researchers who build software to automate our work and foundational AI models that help us see patterns in our data at scale. Together our laboratories, data, software, compute, and team comprise what we believe is one of the most sophisticated operating systems for drug discovery on earth, The Recursion OS.
How we Create Value
We leverage the Recursion OS to deliver value in three ways: 1) our own pipeline of clinical and preclinical potential medicines focused in precision oncology, rare disease, and other niche areas of unmet need; 2) by discovering new medicines with large biopharmaceutical companies in some of the biggest areas of unmet need like neuroscience and inflammation; and 3) by leveraging our tools, technology, and data for the benefit of other partners in targeted and limited ways.
Figure 1. Portfolio poised for value creation from a unified operating system. 1Includes preclinical programs (programs expected to enter the clinic within the next 18 months). 2Program milestones include data readouts, preliminary data updates, regulatory submissions, trial initiation, etc.
10
Table of Contents
2024 Highlights & Progress
In 2024, we accelerated the next wave of AI-driven drug discovery and development, delivering key milestones across multiple clinical programs, advancing our transformative partnerships, unveiling major breakthroughs in foundation models, and by consolidating some of the best tools, technologies, and talent into what we believe is the leading company in the burgeoning field of TechBio.
Advancements in the Pipeline:
•REC-617: A potential best-in-class CDK7 inhibitor optimized using our AI platform, delivered early Phase 1/2 results demonstrating promising safety and preliminary efficacy, including a durable partial response in a late-stage metastatic ovarian cancer patient and stable disease across four other patients with solid tumors (e.g. CRC, NSCLC)
•REC-994: A potential first-in-disease oral superoxide scavenger for symptomatic CCM, confirmed safety and tolerability of chronic dosing in a Phase 2 study, with exploratory analyses suggesting lesion volume reduction on MRI and symptom stabilization as evaluated by change in mRS scores
•Clinical Advancements and Regulatory Milestones: Initiated three clinical studies: DAHLIA (Phase 1/2, REC-1245 for solid tumors and lymphoma), TUPELO (Phase 1b/2, REC-4881 for FAP), and ALDER (Phase 2, REC-3964 for recurrent C. difficile infection), received IND clearance for REC-4539 (small cell lung cancer), CTA approval for REC-3565 (b-cell malignancies), and progressed REC-4209 (idiopathic pulmonary fibrosis) to IND-enabling studies
Advancements in Partnerships:
•Roche and Genentech: Generated whole-genome and chemical perturbation maps in a gastrointestinal oncology indication and a whole genome neuroscience phenomap. The neuro phenomap resulted in the exercise of a $30M milestone
•Sanofi: Achieved $15M in milestones, advancing multiple targets in immunology and oncology into lead optimization
•Bayer: Completed 25 multimodal oncology data packages and delivery of LOWE, our LLM-orchestrated workflow software, to enhance research capabilities
•Merck KGaA (Darmstadt, Germany): Advanced alliance to identify first-in-class or best-in-class targets across oncology and immunology
Advancements in Platform:
•Full Stack AI Powered Platform: Our constantly evolving Recursion OS spans target discovery through clinical development, enabling efficient molecule design and testing for both first-in-class and best-in-class opportunities
•Breakthroughs in Foundation Models: We’ve developed both unimodal and multimodal AI models like Phenom, MolPhenix, and MolGPS that accelerate our ability to make high-confidence predictions in our therapeutics programs
•Advancement in Causal AI Models: Through collaborations with Tempus and Helix, we integrate real-world, scaled patient datasets with our proprietary internal data to deepen biological insights and better match our therapeutic candidates with target populations
•Emerging Focus on ClinTech: We are using AI and machine learning to optimize clinical trial design, accelerate patient enrollment, and enhance evidence generation through data-driven methodologies
What’s Next
In 2025, we plan to accelerate our pipeline with additional clinical readouts, deepen existing partnerships, and further expand our proprietary data suite. We anticipate up to 10 additional clinical program milestones over the next 18 months, exciting milestone achievements in our R&D collaborations, and continued breakthroughs in AI-driven discovery, advancing our mission to decode biology to radically improve lives. Significant milestones from our portfolio of pharma R&D partnerships and a focus on monetizing select assets in our pipeline will help to subsidize continued investment to expand our leading position in the TechBio space, which we view as a generational opportunity for value creation.
11
Table of Contents
Business Overview
Recursion is a leading clinical stage TechBio company with a mission to decode biology to radically improve lives. We aim to achieve our mission by industrializing drug discovery using the Recursion Operating System (OS), a vertical platform of diverse technologies that enables us to map and navigate trillions of biological, chemical, and patient-centric relationships utilizing approximately 65 petabytes of proprietary data.
The Recursion OS integrates ‘Real World’ data generated in our own wet-laboratories or by select partners and a ‘World Model’ which is a collection of AI computational models we also build in-house. Today, our scaled ‘wet-lab’ biology, chemistry, and patient-centric experimental data feed our ‘dry-lab’ computational tools to identify, validate, and translate therapeutic insights, which we can then validate in our wet-lab to both advance drug discovery programs and to generate data to further refine our world model.
Figure 2. The Recursion OS. Recursion generates massive quantities of rich, high-dimensional real world –omics data (e.g., phenomics, transcriptomics and proteomics) and chemical data. We build and train ML and AI models that take the data and learning from the real world to understand and identify patterns and insights, using the fastest supercomputer wholly owned and operated by any pharmaceutical company globally (based on available data). This creates a virtuous cycle of learning and iteration based on real-world data and models that learn to simulate that real world. This virtuous cycle, i.e., the Recursion OS, generates value through our pipeline, in addition to building pipelines for our partners.
We have demonstrated that the Recursion OS already accelerates the timelines and scale of drug discovery, and we hope to prove in the coming years that we not only meet the industry average probability of success in clinical development but exceed it. If successful in achieving our mission, we may be able to build one of the most valuable businesses in our industry while also improving the lives of patients and pushing down the cost of healthcare.
Figure 3. Over time, we believe our transformational way of designing and developing drugs can change the industry's underlying pharmacoeconomic model, what we call 'shifting the curve'. We aim to demonstrate that it is simultaneously possible to improve probability of success through designing better quality drugs while also reducing investment requirements through improved technologies and process. Recursion was created to take advantage of the discontinuity between these fields and harness the power of accelerating technological innovations to improve the efficiency of drug discovery and development.
12
Table of Contents
The Recursion OS – A Platform that Powers a Portfolio
The Recursion OS is a full-stack solution delivering technology-enabled first-in-class and best-in-class molecules with speed, efficiency, and scale from target discovery through early clinical development. We generate and aggregate enormous quantities of high-quality, high-dimensional data spanning hundreds of millions of cellular perturbations across biology and chemistry, translational experiments, ADMET experiments, in vivo experiments, patient data, and from scaled automated chemical synthesis. In parallel, we have built foundation models that leverage those data to learn and understand the underlying biological and chemical interactions with broad predictive capabilities.
Figure 4. Recursion’s World Model approach (1) Profile biological and chemical systems using automation to scale a small number of data-rich assays, including phenomics, transcriptomics, InVivomics, and ADME to generate massive, high quality empirical data; (2) aggregate and analyze the resultant data using a variety of machine learning models, in a process coordinated with in-house software systems and tools; and (3) map and navigate leveraging proprietary software tools to infer properties and relationships in biology and chemistry. These inferred properties and relationships serve as the basis of our ability to predict how to navigate between biological states using chemical or biological perturbations, which we can then validate in our automated laboratories, completing a virtuous cycle of learning and iteration.
Our competitive advantage compared to the hundreds of companies both large and small that seek to follow our path is the intense focus and early success in verticalizing this approach. The company began by pioneering a point solution to scaled target discovery and hit ID based on phenomics, the use of cell morphology and computer vision. Today, we are generating, aggregating, or simulating data from patients to cells, cells to pathways, pathways to proteins, proteins to atomistic interactions, cells to organoids, organoids to animals, and animals to people in the clinic. We're systematically capturing complex, high-dimensional datasets, training specialized machine learning and AI models, and building foundation models that synthesize insights across diverse data layers. As we combine and coalesce these foundation models, spanning target discovery through clinical development, we are increasingly building a ‘world model’ that contains within it a virtual representation of how biology and chemistry are working. Today, our world model is enabling us to make many high-confidence predictions about the result of previously untried or untested questions.
In the coming years we believe that our world model will attain a level of understanding of biology and chemistry of sufficient quality that our wet lab will move from data generation for model improvement as its primary use to ‘scaled validation of simulated solutions.’ In essence, our world model will become a ‘virtual cell’ where we can simulate an inexhaustible quantity of ‘experiments,’ identify those targets and chemistries that have the highest probability of success in modulating disease and achieving a desired (and automatically generated) Target Product Profile, and then our wet lab can validate those predictions at scale.
13
Table of Contents
Figure 5. Over time, models can become broadly applicable and performant enough to be the first rather than last step in the process, that is “move to the left” in the diagram, with Real World Models serving to validate individual insights “on the right.”
Building a Pipeline
Building on a first-in-class and best-in-class OS platform after the business combination with Exscientia, Recursion’s pipeline now encompasses 10 clinical and preclinical programs and over 10 advanced discovery programs across oncology, rare diseases, and other areas of high unmet need. This broad and rapidly evolving portfolio reflects our commitment to advancing discovery and clinical development through unbiased, scaled scientific insights and AI-driven discovery. Programs in our internal pipeline are built on unique biological and chemical insights surfaced through the Recursion OS where:
•The etiology of the disease is well defined, but the subsequent impacts of the disease are generally obscure and/or the primary targets are typically considered undruggable,
•There is a high unmet medical need, no approved therapies, or significant limitations to existing treatments.
Figure 6. The power of our Recursion OS exemplified by our expansive therapeutic pipeline. 1Includes non-small cell lung cancer (NSCLC), colorectal cancer, breast cancer, pancreatic cancer, ovarian cancer, head and neck cancer. 2Joint venture with Rallybio.
14
Table of Contents
Advancing our Pipeline
We are accelerating critical clinical milestones while delivering measurable progress against diseases with high unmet medical needs. At the same time, we continue to validate various components of the Recursion OS, which has played a role in advancing every program in our portfolio, reinforcing its potential to accelerate drug discovery and development.
We have already demonstrated promising safety and preliminary efficacy data for two of our programs. REC-994, a superoxide scavenger in development as a potential first-in-disease therapy for symptomatic CCM, was featured in a late-breaking oral presentation at the 2025 International Stroke Conference. Phase 2 data highlighted MRI-based lesion volume reduction and symptom stabilization trends. Next steps in this program will be informed by regulatory discussions and long-term extension data expected in 2025. REC-617, a potential best-in-class CDK7 inhibitor has demonstrated early clinical activity in advanced solid tumors, including a durable partial response in metastatic ovarian cancer and stable disease across patients with multiple tumor types. These findings support further clinical development as we continue to explore its potential in combination regimens.
In parallel, Recursion has recently initiated three other clinical studies: DAHLIA (Phase 1/2, REC-1245 for solid tumors and lymphoma), TUPELO (Phase 1b/2, REC-4881 for FAP), and ALDER (Phase 2, REC-3964 for prevention of recurrent C. difficile infection). Additionally, REC-3565 (B-cell malignancies) and REC-4539 (SCLC) are expected to enter dose escalation studies. REC-4209 (idiopathic pulmonary fibrosis) has progressed to IND-enabling studies and REV102 (HPP) IND-enabling studies have been initiated, further expanding our pipeline across oncology, rare diseases, and other areas of high unmet need.
Anticipated Near-term Catalysts. Recursion is poised for a catalyst-rich period, with multiple programs reaching critical milestones over the next 18 months. REC-3565 (MALT-1i) will enter Phase 1 for B-cell malignancies, with the first patient expected to be dosed in the first half of 2025. REC-617 (CDK7i) will initiate combination studies in advanced solid tumors in the first half of 2025. REC-4881 (MEK1/2i) will report Phase 1b/2 safety and early efficacy data for FAP during the same period. REC-2282 (HDACi) for NF2-related meningioma will undergo PFS6 futility analysis and REC-4539 (LSD-1i) will begin Phase 1 dose escalation in SCLC during the same period. Additional Phase 1 data from the ELUCIDATE trial for REC-617 (CDK7i) in advanced solid tumors is expected in the second half of 2025. Furthermore, we continue to rapidly advance development programs such as a PI3Kα H1047Ri and an ENPP1i towards the clinic.
Impact Through Partnerships
Through our business combination with Exscientia, we have doubled our partnership footprint with leading pharmaceutical companies including Roche and Genentech, Sanofi, Bayer and Merck KGaA (Darmstadt, Germany), securing $450M million in upfront milestone payments to date with the potential for over $20 billion in additional milestones before royalties. These global collaborations not only provide near-term cash flows but also combine our scaled biology, precision chemistry, and automated synthesis capabilities to pave the way for transformative therapies in oncology, neuroscience, immunology, and other therapeutic areas with high unmet need. By partnering with some of the best biopharmaceutical companies on earth in their respective areas, our platform and team have an opportunity to learn from some of the most experienced teams in the industry. By uniting our AI-driven platforms, vast proprietary data, and deep scientific expertise, we continue to unlock powerful innovations and expand patient impact. Below are some of the latest developments illustrating this momentum:
Roche and Genentech
•Gastrointestinal-Oncology Advancements: In partnership with Roche and Genentech, we generated multiple whole-genome phenomaps with chemical perturbations across various disease-relevant cell types, enabling deeper insights into how different cellular contexts respond to gene knockouts and chemicals.
•Neuro-specific CRISPR KO Phenomap: In partnership with Roche and Genentech, we have developed the first whole-genome CRISPR knockout map in neural iPSC cells, providing valuable data to identify potential new targets in neuroscience, an area with limited new discoveries.
•Milestones and Collaboration: The neuroscience phenomap work led to a $30M option from Roche and Genentech in August 2024, and we’re moving forward with target validation projects.
Sanofi
•Immunology & Oncology Achievements: We reached milestones in three programs, generating $15M in aggregate payments from Sanofi for two of these programs in 2024.
Bayer
•Oncology Achievements: Completed 25 multimodal oncology data packages utilizing the Recursion OS platform. Multiple programs rapidly progressing to Lead Series nomination.
•LOWE: Additionally, Bayer was the first beta-user of our LOWE LLM-orchestrated workflow software to enhance their research capabilities.
15
Table of Contents
Merck KGaA (Darmstadt, Germany)
•Our ongoing alliance with Merck KGaA, Darmstadt, Germany is focused on leveraging Recursion’s discovery engine to identify first-in-class and best-in-class targets across oncology and immunology, driving innovation in these key therapeutic areas.
Leading indications of success in our business combination with Exscientia
We believe that to truly redefine the way drugs are discovered and developed, we need to build technology tools and data across the full-stack of drug discovery and development. Doing this well also requires the best talent from across many different fields. While Recursion has, in our view, achieved more than any other TechBio company, we constantly survey the market to find potential opportunities to augment and accelerate our business. The business combination with Exscientia was one such opportunity, bringing together the best biology-first TechBio platform in Recursion and one of the most comprehensive chemistry-first TechBio platforms in Exscientia, a compelling set of both first and best-in-class clinical programs, sector-leading partnerships and some of the best talent in the industry.
Given the compressed timeline between signing and close of the transaction (just over 3 months), prior to the close we prioritized in-depth assessments of each company's programs, partnerships, technology, capabilities, and ways of working. With this information in hand, we then conducted an organizational design process that allowed us to provide go-forward status and role clarity for each employee within 48 hours of close. We also outlined a set of goals to be accomplished in the first 90 days post close that will help us rapidly demonstrate the exponential value of the business combination and help employees to rapidly integrate across teams, portfolio, partnerships, and technologies. We are already seeing the benefits of this approach in terms of amplifying delivery, though much more work remains to be completed through 2025.
In the weeks post-close we focused on ensuring all employees understood where they were situated within the go-forward organization (role, manager, etc.) and connected them with new members of their teams. We wrapped 2024 with a 2-day Welcome Event in London, giving former Exscientians the opportunity to learn about Recursion, meet new colleagues across a variety of teams, celebrate their contributions to the deal close, and begin to shift their professional identity to that of a Recursionaut.
90 Day Goals
Before the transaction closed, we established a set of key goals to be achieved within approximately 90 days. These goals were designed to quickly demonstrate early indicators of the value of combining the companies. The goals largely focus on deploying relatively mature and unique elements of each company's technology to accelerate the other party's programs and processes. These goals also require key members of each legacy team to sprint together, accelerating the team forming and norming process. Key updates are as follows:
Pipeline
Using Recursion’s causal AI to optimize Exscientia clinical programs - LSD1 Patient Stratification:
•Recursion is utilizing AI models and Tempus data to build a patient stratification framework in small cell lung cancer (SCLC). This work is informing clinical strategies for the planned REC-4539 Phase 1 study that originated at Exscientia and is commencing in the first half of 2025.
•We have expanded this work to explore indications for REC-4539 beyond SCLC and laboratory validation work is beginning imminently.
•This work will be used as a template to expand causal modeling for many programs beyond LSD1.
Use Exscientia’s Centaur Platform to Accelerate Recursion’s Internal Programs:
•Programs from legacy Recursion in early chemistry design cycles have already entered Exscientia’s Centaur precision chemistry platform, where significant improvements in potency have been demonstrated.
•Advanced protein structure predictions are guiding compound optimization, aiming to enhance binding conformation and optimize key properties.
•In-progress synthesis and cryo-EM work are enhancing our understanding of binding interactions, informing the next design cycles and optimizing compound characteristics.
Partnerships
Use Recursion’s Maps of Biology to Identify New Targets for Legacy Exscientia Partners:
•The Recursion OS has been used to identify hit compounds in 7 immune-relevant targets or dual target pairs and early validation work has commenced to prepare reports for our partners.
16
Table of Contents
Accelerate Recursion’s Partnered Programs Leveraging Exscientia’s Centaur Chemistry Platform:
•We integrated Centaur into more than 10 design cycles for programs Recursion has previously partnered, with early validation work achieved and progress accelerating across multiple additional partnered programs.
•We have successfully delivered compounds for wet-lab testing on Recursion’s platform using the legacy Exscientia automated synthesis platform for partnered programs.
•Applied a newly built pipeline using structural bioinformatics and molecular dynamics for more precise compound design, focused on addressing specific design challenges on partnered programs.
Platform
Incorporate Exscientia Tech into Recursion’s Workflows:
•Reduced manual effort by 60% for evidence collection for hit nomination packages supporting entry into hit-to-lead, through knowledge graphs and LLM-based data aggregation with further reduction expected with additional data layers.
•Mapped 1.4 million active ligands to binding pockets for structure-based drug discovery and target deconvolution.
•Working to leverage Exscientia’s tools to achieve a 75% reduction in time-to-program nomination by increasing alignment with portfolio strategy and bridge phenotypic- to target-based drug discovery and facilitating target-centric and structure-based drug discovery within Centaur Chemist.
Integrate Recursion models into Centaur:
•To augment Exscientia’s chemical design platform with additional filters, experimental data for >950,000 compounds profiled in living cells were used to build two new models for (1) measurable cellular activity and (2) cytotoxicity. These large datasets provide highly generalizable models demonstrating a >2.5x increased efficiency in detecting new bioactive scaffolds with a >40% reduction in the flagging of likely cytotoxic compounds.
•18 new combined ADMET data products and respective models representing the combined organization’s datasets, and five program-specific phenomics response models have been added to Centaur, enabling use of legacy Exscientia’s active learning based precision design technology with Recursion internal programs.
Corporate Updates
•In order to continue to streamline its operations and focus on its core geographic footprint, the Company's fully owned subsidiary, Exscientia AI Limited ("EXAI"), will carve out its Austrian operations. On February 27, 2025, the Company entered into an agreement to invest in and acquire 49% of the share capital of Alpha Biotechnology GmbH (“Alpha”), an Austrian startup that will leverage a patient-tissue platform for the development of precision therapeutics for the treatment of hematological and solid cancers, while focusing its efforts and moderating spend. Alpha will in turn acquire 100% of the share capital of Exscientia GmbH, a wholly-owned Austrian subsidiary of EXAI which holds certain intellectual property assets crucial to Alpha’s operations and business activities. The closing of the transaction is expected in the first half of 2025, subject to customary closing conditions.
A clear vision, durable mission and a focus on people and culture drive success
Recursion was founded in 2013 with a vision to capitalize on the convergence of advancements in computation and machine learning to address the decreasing efficiency of drug discovery and development. We believe that this opportunity represents one of the most positively impactful applications of ML and AI. Our vision is to leverage technology to map and navigate biology, chemistry, and patient-centric outcomes to increasingly transition the process of developing medicines from discovery to design. We believe that advanced computational approaches, massive datasets or human intelligence alone cannot fundamentally shift the efficiency curve of drug discovery and development; instead, we believe that those companies that augment their teams with sophisticated computational tools and focus deeply on generating and aggregating the right datasets will have a significant advantage. Our success and the success of the burgeoning TechBio sector has the promise to drive more, new, and better medicines to patients at higher scale and lower prices in the coming decades. We are working to not only lead this space – but define it.
Our mission at Recursion, Decoding Biology to Radically Improve Lives, flows naturally from our vision. We interpret our mission expansively and believe it to be a durable direction and source of inspiration for our team. We seek not only to radically improve the lives of patients who could benefit from the medicines we help to deliver, but the lives of those who care for those patients, the lives of our employees and their families, as well as the communities in which we operate our company.
We’ve intentionally designed our culture to fuel the pursuit of our mission. Our Founding Principles are guideposts for scientific and technical decisions and our Values underpin how our employees engage day-to-day with colleagues inside and outside the company. The Recursion Mindset, a deep commitment to achieving impact at unprecedented scale through new industrialized approaches, is an essential component of building our TechBio ecosystem. Our employees bring all these to life, contributing their unique expertise and experiences from their incredible breadth of fields and industries. For all of our employees, Recursion is a unique company with a different way of working.
17
Table of Contents
Figure 7. Recursion’s team requires operating at the interface of many diverse fields. Building a TechBio company requires fluency in operating at the interface of many disciplines and fields not previously attuned to working as closely in traditional biopharma.
18
Table of Contents
Figure 8. Recursion’s Founding Principles and Values support our ambitious mission. Together, these elements shape Recursion’s culture by guiding our people to high-impact decision-making and behaviors.
19
Table of Contents
Recursion In-Depth
TechBio: The Industrialization of Drug Discovery and Development Problem
The traditional drug discovery and development process is characterized by substantial financial risks, with increasing and long-term capital outlays for development programs that often fail to reach patients as marketed products. Historically, it has taken over ten years and an average capitalized R&D cost of approximately $2 billion per approved medicine to move a drug discovery project from early discovery to an approved therapeutic. Such productivity outcomes have culminated in a rapidly declining internal rate of return for the biopharma industry.
Figure 9. Historical biopharma industry R&D metrics. The primary driver of the cost to discover and develop a new medicine is clinical failure. Less than 4% of drug discovery programs that are initiated result in an approved therapeutic, resulting in a risk-adjusted cost of approximately $2.3 billion per new drug launched.1,2,3,4,5
Despite significant investment and brilliant scientists, these metrics point to the need to evolve a more efficient drug discovery process and explore new tools. Traditional drug discovery relies on basic research discoveries from the scientific community to elucidate disease-relevant pathways and targets to interrogate. Coupled with biology’s incredible complexity, this approach has forced the industry to rely on reductionist hypotheses of the critical drivers of complex diseases, which can create a ‘herd mentality’ as multiple parties chase a limited number of therapeutic targets. The situation has been exacerbated by human bias (e.g., confirmation bias and sunk-cost fallacy). Accentuating this problem, the sequential nature of current drug discovery activities and the challenges with aggregation and interoperability of data across projects, teams and departments lead to frequent replication of work and long timelines to discharge the scientific risk of such hypotheses. Despite decades of accumulated knowledge, the result is that drug discovery has unintentionally created hurdles for innovation.
Simultaneously, exponential improvements in computational speed and reductions in data storage costs driven by the technology industry, coupled with the rapid rise of LLMs, generative AI and other ML tools, have transformed complex industries from media to transportation to e-commerce. Historically, the biopharma sector has been slow to embrace such innovations. Over the past 2-3 years, there have been remarkable shifts in perception among technology and biopharma companies as well as among regulators and policymakers, who highlight the utility of AI/ML for broad drug discovery and development from novel target discovery to automated chemistry synthesis and next-generation manufacturing. We believe this rapid acceleration and adoption of these technologies demonstrates the growing consensus that AI/ML is a catalyst for substantial leaps in drug discovery.
1 Zhou, S. and Johnson, R. (2018). Pharmaceutical Probability of Success. Alacrita Consulting, 1-42
2 Steedman M, and Taylor K. (2024). Measuring the return from pharmaceutical innovation. Deloitte. 1-28.
3 DiMasi et al. (2016). Innovation in the pharmaceutical industry: New estimates of R&D costs. Journal of Health Economics. 47, 20-33.
4 Paul, et al. (2010). How to improve R&D productivity: the pharmaceutical industry’s grand challenge. Nature Reviews Drug Discovery. 9,203-214
5 Martin et al. (2017). Clinical trial cycle times continue to increase despite industry efforts. Nature Reviews Drug Discovery. 16, 157
20
Table of Contents
Opportunity
Late-stage clinical failures are the primary driver of reduced impact and IRR in today’s pharmaceutical R&D model. Reducing the rate of costly, late-stage failures would be the most compelling way to achieve a more productive drug discovery and development process, though expanding the areas of biology that can be explored, accelerating the timeline from hit to a clinical candidate, decreasing the costs of discovery and creating more scalable systems would also create a more sustainable R&D model, all else held equal. To achieve this more sustainable model, we believe that in its ideal state, a drug discovery funnel would morph from the being shaped like the letter ‘V’ to being shaped like the letter ‘T,’ where a broad set of possible therapeutics could be narrowed rapidly to the best candidate in a scaled and efficient way. Subsequently, programs that advance through the remaining steps of the discovery and development process would proceed quickly and with no attrition. While such a path is impossible to fully achieve, rapidly improving technology tools across biology, chemistry and computation are creating the conditions where, in the right hands, progress towards this ‘T’-shaped funnel is possible.
Figure 10. Reshaping the drug discovery funnel. Recursion’s goal is to leverage technology to reshape the typical drug discovery funnel towards its ideal state by moving failure as early as possible to rapidly narrowing the funnel into programs with the highest probability of success.6,7,8,9
The Recursion OS provides an opportunity for mapping and navigating massive biological and chemical datasets that contain trillions of inferred relationships between disease-causative perturbations and potentially therapeutic compounds. Collectively, the components of the Recursion OS can be joined together in a modular way to identify, validate and advance a broad portfolio of novel therapeutic programs quickly, cost-effectively and with minimal human intervention and bias - industrializing drug discovery and development. We use standardized, automated workflows to identify programs and advance them through key stages of the drug discovery and development process which includes:
•Patient Connectivity and Novelty (i.e., program initiation)
•Hit and Target Validation
•Compound Optimization
•Translation
•IND-Enabling Studies
•Clinical Development
6 Steedman M, and Taylor K. (2020). Ten years on: Measuring the return from pharmaceutical innovation. Deloitte. 1-44.
7 DiMasi et al. (2016). Innovation in the pharmaceutical industry: New estimates of R&D costs. Journal of Health Economics. 47, 20-33.
8 Paul, et al. (2010). How to improve R&D productivity: the pharmaceutical industry’s grand challenge. Nature Reviews Drug Discovery. 9,203-214
9 Martin et al. (2017). Clinical trial cycle times continue to increase despite industry efforts. Nature Reviews Drug Discovery. 16, 157
21
Table of Contents
Figure 11. Multiple -omics modalities within Recursion OS form a full-stack platform that spans the drug development pathway, incorporating patient-centric scaled biology, target exploration, hit discovery, lead optimization, precision chemistry design, automated chemical synthesis, predictive ADMET (absorption, distribution, metabolism, excretion, and toxicity), biomarker selection, translational capabilities, and clinical development. Areas from legacy Exscientia platform indicated in orange.
We believe we have made progress in reshaping the traditional drug discovery funnel in the following ways:
•Broaden the funnel of therapeutic starting points. Our flexible and scalable mapping tools and infrastructure enable us to infer trillions of relationships between human cellular disease models and therapeutic candidates based on real empirical data from our own wet labs.
•Identify failures earlier when they are relatively inexpensive. Our proprietary navigation tools enable us to explore our massive biological, chemical, and patient-centric datasets to validate more and varied hypotheses rapidly. While this strategy results in an increase in early-stage attrition, the system is designed to rapidly prioritize programs with a higher likelihood of downstream success because they have been explored in the context of high-dimensional, systems-biology data. Over time, and as our OS improves, we expect that moving failure earlier in the pipeline will result in an overall lower cost of drug development.
•Optimize molecular design and synthesis through Centaur Chemist. Our Centaur Chemist platform integrates AI-driven generative design with automated synthesis, enabling rapid iteration and optimization of new chemical entities. By leveraging predictive modeling for potency, selectivity, and ADMET properties, we can efficiently generate high-quality, differentiated drug candidates while reducing synthesis timelines and experimental bottlenecks.
•Accelerate delivery of high-potential drug candidates to the clinic. The Recursion OS contains chemistry tools that enable highly efficient exploration of chemical space as well as translational tools that improve the robustness and utility of in vivo studies.
•Enhance clinical development efficiency through ClinTech. We are applying machine learning and AI to optimize clinical trial design, accelerate patient enrollment, and enhance evidence generation. By integrating scaled patient data with predictive analytics, we aim to improve patient stratification, match therapies with the right populations, and reduce trial failure rates—advancing high-potential medicines to patients faster and more efficiently.
By leveraging our Recursion OS to explore and advance our programs, we have shown leading indicators of improvement when compared to the traditional drug discovery process, particularly with respect to cost and time. We believe that combining Exscientia’s state of the art chemical platform with Recursion’s cutting edge OS will further enable us to (i) identify candidate compounds earlier in the research cycle, (ii) spend less per program, and (iii) expedite our drug discovery progress compared to industry. Across >280 Recursion programs from late 2017 through 2024 the average amount of time to reach the validated lead stage is less than 13 months. We also use AI/ML tools to better understand which molecules to make and test and ultimately design better quality molecules that can solve complex problems – on average, the industry synthesizes about 2,500 molecules to candidate, both legacy Recursion and legacy Exscientia synthesized 250, respectively.
22
Table of Contents
Ultimately, we believe that future iterations of the Recursion OS will enable even greater improvements minimizing the total dollar-weighted failure and maximizing the likelihood of success allowing us to deliver better quality medicines to more patients in need.
Figure 12. (Far Left): Time from hypothesis screening to validated hit package for legacy Recursion programs. (Center Left): Legacy Exscientia compounds synthesized from hit to candidate ID. (Center Right): Total spend from hypothesis screening to the completion of IND-enabling studies for legacy Recursion novel chemical entity (NCE) programs that advanced to clinical trials. The cost to IND has been inflation-adjusted using the US Consumer Price Index (CPI) (Far Right). Time to validated lead is the average of >280 legacy Recursion programs since late 2017 through 2024.10
The Recursion OS has not only improved speed and cost but also led us to explore novel targets which could give us a competitive advantage where multiple parties often simultaneously pursue a limited number of similar target hypotheses. Below one can see quantitative measures for how we prioritize programs characterized by (i) strong genetically driven biological evidence and (ii) differentiated novel biology.
10 Paul, et al. (2010). How to improve R&D productivity: the pharmaceutical industry’s grand challenge. Nature Reviews Drug Discovery. 9,203-214
23
Table of Contents
Figure 13. We use LLMs and software tools to organize and initiate internal programs using both proprietary and public data. We prioritize programs at scale by focusing on targets where our proprietary data provides a distinct arbitrage that suggests we can drive towards novel target identification and selection in oncology. Each circle represents a gene that can be searched by the Recursion OS across several biological and pharmacological factors. LLMs harness public datasets such as Cancer Dependency Map, Open Targets, TCGA, CCLE, and COSMIC and Recursion proprietary datasets such as phenomap inferences, Matchmaker assessments, InVivomics experiments, and ADME predictions.11
Approach
At Recursion, we are pioneering the integration of innovations across biology, chemistry, automation, data science and engineering to industrialize drug discovery in a full-stack solution across dozens of key workflows and processes critical in discovering and developing a drug. For example, by combining advances in high content microscopy with arrayed CRISPR genome editing techniques, we can rigorously profile massive, high-dimensional biological and chemical perturbation libraries in multiple human cellular contexts to create digital ‘maps’ of human biology. Leveraging advances in scaled computation, we can conduct massive virtual screens to predict the protein targets for billions of chemical compounds. Similarly, data generated from our automated DMPK module and InVivomics platform enables us to predict ADMET properties and identify toxicity signals, respectively, significantly faster than traditional methods. And now, with the business combination with Exscientia complete, we can drive many of our programs from hit to development candidate using an automated internal chemical synthesis platform.
11 Ochoa, et al. (2023). The next-generation Open Targets Platform: reimagined, redesigned, rebuilt. Nucleic acids research. 51(D1): D1353-D1359
24
Table of Contents
Figure 14. Recursion’s approach to drug discovery. We utilize our Founding Principles on the right to build datasets which are scalable, reliable and relatable in order to elucidate novel biological and chemical insights and industrialize the drug discovery process.
We have used our approach to generate, aggregate, and integrate one of the largest proprietary biological, chemical, and patient-centric datasets in the world at approximately 65 petabytes at the end of 2024. This dataset includes proprietary phenomics, transcriptomics, predicted protein-ligand binding interactions, InVivomics, ADMET data, and more across many biological and chemical contexts as well as preferred access to over 20 petabytes of multimodal oncology patient data from Tempus. Additionally, we have built a proprietary suite of software applications within the Recursion OS which has identified over 7 trillion predicted biological and chemical relationships. With our approach, we endeavor to turn drug discovery into a search problem where we map and navigate biology in an unbiased manner to discover new insights and translate them into potential new medicines at scale.
Competitive Landscape and Differentiation
There are a few key factors that differentiate Recursion from other technology-enabled drug discovery companies.
1.Recursion has built a full-stack platform utilizing many biology, chemistry, and patient-centric proprietary datasets and modular tools to industrialize drug discovery, while most other competitor companies rely on a point solution to solve one important step in drug discovery. We recognize that drug discovery is made up of many steps, and a point solution is insufficient to generate efficiencies across the entire process. To decode biology, we must construct a full-stack technology platform capable of integrating and industrializing many complex workflows.
2.Recursion integrates wet-lab and dry-lab capabilities in-house to create a virtuous cycle of iteration. Fit-for-purpose wet-lab experimental data are translated by dry-lab digital tools into in silico hypotheses and testable predictions, which in turn generates more wet-lab data from which improved predictions can be made. Recursion is well positioned compared to companies of a similar stage either focused more specifically on the wet-lab only (traditional biotech or pharma companies) or dry-lab only (companies facing rapidly commoditized algorithms and a challenge differentiating on non-proprietary data).
3.Recursion has achieved a significant scale with respect to its scientific, technological, and business endeavors. With eight clinical-stage programs, an exciting preclinical pipeline, four of the largest discovery partnerships in the biopharma industry with Roche/Genentech, Sanofi, Bayer and Merck KGaA (Darmstadt, Germany), and four technology-focused partnerships, Recursion has achieved a scale, level of integration, and stage that few other TechBio companies have.
25
Table of Contents
Value Drivers
While most small to medium-sized biopharma companies are focused on a narrow slice of biology or a single therapeutic area, the Recursion OS allows us to discover and translate at scale across biology. However, we are cognizant that building disease-area expertise, especially in clinical development, is essential. We have developed a multi-pronged, capital-efficient business model focused on key value-drivers that enable us to demonstrate our progress over time while continuing to invest in the development of the Recursion OS, which we believe is the engine of value creation in the long-term. While our mapping, navigating and designing tools have the plasticity to be applied across therapeutic areas and modalities, our business model is tailored to maximize value and advance programs cost-effectively based on the nature of market and regulatory dynamics associated with our three value drivers (internal pipeline, transformational partnerships, and fit-for-purpose proprietary biological, chemical, and patient-centric data).
Figure 15. We harness the value and scale of our Recursion OS using a capital efficient business strategy. Our business strategy is segmented into our following value-drivers: (1) internally developed programs in capital-efficient therapeutic areas; (2) partnered programs in resource-intensive therapeutic areas; and (3) proprietary, fit-for-purpose data and models.
Value Driver 1 - Internally Developed Programs in Capital Efficient Therapeutic Areas
We believe that the primary currency of any biotechnology company today is clinical-stage assets. These programs can be valued using a variety of models by stakeholders in the biopharma ecosystem and most importantly, present the potential to meet critical patient needs. For Recursion, these assets have a variety of additional benefits, including: (i) validation of key elements of the Recursion OS, (ii) growing our expertise in clinical development and (iii) building in-house processes to facilitate smooth interaction with regulatory agencies and advance medicines towards the market. If the Recursion OS evolves as designed, then it will continuously improve with more iterations such that future programs could be more novel and potentially more valuable than today’s programs. Operating as a vertically integrated TechBio company that leverages technology at every step from target discovery through clinical development (and even marketing and distribution) may be the long-term business model with the most upside for our stakeholders, including both investors and patients. We may be opportunistic about selling or licensing assets after they achieve key value-inflection milestones so that we can re-invest in our long-term strategy.
Value Driver 2 - Partnered Programs in Resource Intensive Therapeutic Areas
We believe that in its current form, our Recursion OS is already capable of delivering many more therapeutic insights than we would be able to shepherd alone today. As such, we have chosen to partner with experienced, top-tier biopharma companies like Roche and Genentech, Sanofi, Bayer, and Merck KGaA (Darmstadt, Germany) to explore intractable and resource-intensive areas of biology. The key advantages of these partnerships are that: (i) we are able to deploy the Recursion OS to turn latent value into tangible value in areas of biology where it would be challenging for us to do so alone; (ii) the clinical development paths for these large therapeutic areas are often resource-intensive and highly complex; and (iii) we are able to learn from our colleagues at these top-tier companies such that it could give us a competitive advantage in the industry over the longer term. This strategy also embeds us in the discovery process of large pharmaceutical companies and gives rise to an alternative long-term business model whereby we become a valued partner of many such companies. Based on how value is ascribed across our industry today, this model alone is not yet feasible to maximize our business impact. However, we feel that due to shifts within the biopharmaceutical industry there is some potential for this portion of our business model to accrete notable value over the long-term.
26
Table of Contents
Value Driver 3 - Proprietary, Fit-for-Purpose Training Data and Models
As has been demonstrated in many other industries, a value driver and competitive advantage can be generated from the creation of a proprietary dataset. At Recursion, we have generated what we believe to be one of the largest fit-for-purpose, relatable biological, chemical, and patient-centric datasets on Earth. Spanning multiple omics technologies and more than 300 million unique experiments, the approximately 65 petabytes of data that Recursion generates, aggregates, and integrates has the fundamental purpose of being used to train machine learning models. Through intensive internal work, Recursion uses this data and our own models, algorithms, and software to advance our own internal pipeline of medicines (Value Driver 1) as well as in partnership with our collaborators to advance additional discovery programs (Value Driver 2). As our field increasingly recognizes the potential for a technology-driven revolution in drug discovery, our data has increasing potential to drive value directly. We increasingly see the potential to license select models and subsets of our data to a growing universe of collaborators for which internal efforts would be minimal, but value could be significant.
A Platform to Industrialize Drug Discovery
We have generated one of the largest relatable data sets in biopharma using our automated high throughput labs, which run over 2 million experiments per week. Our data includes cellular phenomics, captured using Brightfield microscopy, as well as chemical synthesis, transcriptomics, proteomics, ADMET, InVivomics, genomics and patient data. In total, we have approximately 65 petabytes of proprietary data which we use to train our algorithms and build our Maps of Biology.
In our relentless drive to continue building the most advanced full-stack AI-enabled discovery and development platform, we have now integrated Exscientia's Centaur Chemist platform for molecule design and automated synthesis capabilities with the Recursion OS, allowing us to rapidly move from target discovery to in silico design to physical compound testing.
And while we work daily to continue solidifying our data moat through wet-lab experimentation and simulation, our partnerships with Helix and Tempus give us access to hundreds of thousands of patient insights – including whole exome and whole genome sequencing – across a wide range of chronic diseases and in oncology. By integrating even limited patient data into the Recursion OS, we can derive powerful new insights that directly fuel our pipeline. We’re also using patient data to match our drugs to the specific patient population most likely to benefit, to improve the probability of success of clinical trials – where 90% of drugs in development fail.
Our unique platform approach has continued to evolve over time and we continue to lead the industry in innovation and delivery of potential treatments through our pipeline and partnerships. When we first developed our phenomics-based biological map-making methods using HUVEC cells – creating over 100 billion cells per year for high throughput experiments, our work was dismissed by many. Now, the early success of our pipeline and partnerships, and the broad adoption of our phenomics approaches across nearly every large pharma company in the industry, suggests a much deeper impact. While others onboard technology we pioneered over a decade ago, we have moved to live-cell brightfield imaging and in partnership with Roche and Genentech, we built specific cell manufacturing technologies that derive neurons from hiPSCs at scale – ultimately producing over 1 trillion hiPSC-derived neuronal cells to build the world’s first whole-genome neuronal phenotype or “Neuromap,” triggering a $30 million payment.
Through our unique dataset and compute power, in the past year, we’ve launched a number of breakthroughs in foundation models – including powerful multimodal models like Phenom, MolPhenix and MolGPS. These models give Recursion deeper insights into underlying disease biology and how cells might respond to treatment with new drug candidates, and provide the company with a distinct advantage when driving decisions about which therapeutic programs to pursue.
The Recursion OS
The Recursion OS is the integrated technical and scientific vertical platform that underpins drug discovery and development at Recursion, from program initiation through our clinical trials. Collectively, the components of the Recursion OS can be joined together in a modular way to identify, validate, and advance a broad portfolio of novel therapeutic programs quickly, cost-effectively, and with minimal human intervention and bias. Connecting modules into a system allows us to quantitatively measure impact and make transformative improvements not just within local point solutions, but across the end-to-end drug pipeline, today and into the future.
To achieve this pipeline impact, the modules of the Recursion OS are connected by industrialized workflows that have been standardized, scaled, and automated. To drive greater efficiency, we approach the building of modules and workflows similar to modular programming, but for biology and chemistry, so that the same fundamental capabilities are transferable across drug discovery and development activities and reflected in a diverse portfolio for both our internal pipeline and large pharma partnerships.
27
Table of Contents
In 2024, we:
•Increased the sophistication of modules that improve earlier parts of the pipeline by integrating chemistry-centric models from Exscientia, including our industrialized workflows for chemical optimization enabling both functional- and target-based discovery, biology- and chemistry-centric approaches, and first-in-class and best-in-class compound opportunities.
•Integrated causal models and other analytics based on real patient data into both program initiation phases, ensuring patient connectivity and novelty, as well as clinical development activities, including patient stratification.
A unique advantage of this modular approach is that assay and model outputs become a long-lasting data asset. This consistency, reliability and standardization allows us to build petabytes of data that can be connected across biological, chemical, and patient-centric sources, and across years of experiment execution, model insights, and data types.
Each type of data that Recursion generates in this manner becomes a vertical layer: a computable dataset that can be used to build increasingly ever more complete and generalizable machine learning models to infer biological and chemical states (properties) and relationships, iteratively mapping and navigating trillions of inferred relationships between disease-causative perturbations and potentially therapeutic compounds.
Figure 16. Recursion’s World Model approach (1) Profile biological and chemical systems using automation to scale a small number of data-rich assays, including phenomics, transcriptomics, InVivomics, and ADMET to generate massive, high quality empirical data; (2) aggregate and analyze the resultant data using a variety of machine learning models, in a process coordinated with in-house software systems and tools; and (3) map and navigate leveraging proprietary software tools to infer properties and relationships in biology and chemistry. These inferred properties and relationships serve as the basis of our ability to predict how to navigate between biological states using chemical or biological perturbations, which we can then validate in our automated laboratories, completing a virtuous cycle of learning and iteration.
The creation of virtuous cycles of physical experiments and in silico models has been a competitive advantage for leaders in many industries outside of biopharma. In drug discovery, virtuous cycles of experimentation and machine learning predictions is an approach to efficiently map and navigate biology and chemistry at unparalleled scale and efficiency. Critically, at the scale Recursion operates, whole systems can be systematically and experimentally evaluated, such as the cellular level effect of individual gene knockouts not just for individual genes of interest, but across the whole genome. Such comprehensive Real World data sets can underpin new generation AI-enabled Model World predicted states, where rather than individual model predictions, we create model representations to reason about and prioritize opportunities across large swaths of biology or chemistry, such as the hundreds of thousands of protein-protein interactions across the human interactome.
Traversing across layers can allow us to extract more insights and higher confidence than any layer on its own. For example, patient data is the most relevant to human health and disease of all data modalities, but it suffers deeply from being intrinsically noisy, incomplete, expensive and difficult to collect. In comparison, cellular phenomics data can be collected cheaply, at scale, with extreme data reliability and completeness.
In 2024, we have demonstrated that by combining patient level data with cellular level data, we can extract genetic causal targets from small (24,000) patient data sets that had previously required patient data sets of over one million in some cases to overcome challenges with patient data quality. Similarly, we are exploring how protein-level data can add pathway interpretability and completeness to our cellular level data. We believe that by computationally combining data layers across the micro-meso-macro, we can unlock many new, powerful efficiencies and insights along the drug discovery pipeline.
28
Table of Contents
Figure 17. Highlights of Recursion’s achievements to-date in generating real world data and building a world model demonstrate progress on our delivery of Recursion’s mission to decode biology.12,13
Towards a Virtual Cell – How Recursion is combining data and compute across different scales
By closely integrating real world experimentation and AI in an iterative manner, and across multiple ‘levels’ of biology, one can create cycles of virtuous learning, where large fit-for-purpose wet-lab datasets support better in silico model generation and enable more focused future wet-lab experiments. Over time, this allows something powerful to emerge: rather than a data-first approach underpinning World Model creation, our drug discovery opportunities emerge from the World Model, and Real World physical experimentation serves to validate the most promising insights. The order of operations has swapped from experiments underpinning the creation of models, to a more scientifically comprehensive yet still more efficient approach: World Model predictions being selectively confirmed by Real World experiments. In essence, we will have created a Virtual Cell which we can test in nearly unlimited ways, selecting the most promising outcomes for validation in a Real World cellular system.
Figure 18. Over time, models can become broadly applicable and performant enough to be the first rather than a step in the process, that is “move to the left” in the diagram, with Real World Models serving to validate individual insights “on the right”
12 Internal data and analysis (2025).
13 TOP500 List (2024). https://top500.org/lists/top500/list/2024/06/
29
Table of Contents
Because the utility and impact on drug discovery would be so profound if achieved, a handful of compelling organizations, spanning academics to institutions to companies, are competing to build a Virtual Cell. Success in this endeavor requires scaled hiqh-quality data, spanning multiple levels of biology alongside cutting-edge compute approaches. We believe Recursion is uniquely positioned to lead at the intersection of these needs.
Scale and Quality. Scale is achieved through three pillars:
•Automation: Automation powers our labs, ensuring we create high-quality data outputs at industry leading scale.
•Relatability: Output data is highly relatable allowing us to infer relationships and identify connections across many experiments.
•In silico models: Scaled relatable data is used to train in-silico models to infer and predict experimental outcomes at scale far exceeding what is possible in the real world, which can then be validated at scale in our automated labs.
Spanning multiple layers of biology
Our Real World data layer and World model are built to use data across three scales: macro, meso and micro. The following section provides examples of these models based on each scale.
•Macro: At the highest level, macroscale data informs on organism-level phenotypes and is typically deeply associative rather than causal – one can find associations between clinical outcomes and human observables, but the causal chain of how to go from variant to macro-scale phenotype is usually hidden. Different size scales and data generation formats tend to have differing levels of interpretability, data quality and noise, costs, and data completeness.
•Meso: In the middle, mesoscale data is measurements on cellular biological systems, such as phenomics and transcriptomics. Here, we built our first in silico maps of biology and have historically executed our largest laboratory experiments.
•Micro: At the smallest end, microscale data informs on molecular-level events like protein-small molecule binding and interactions and can enable insights at the level of target- and protein-interactions and properties. It is the realm of many of our physics- and chemistry-centric models, and our protein sciences and biochemical laboratory assays.
Figure 19. Formation of a highly predictive Virtual Cell from scales of data layers that form Real World data and World model. The macro data layer aims to find associations between clinical outcomes and human observables and determine causal chains of biology. The mesocale data uses measurements on cellular biological systems, such as phenomics and transcriptomics, to build in silico maps of biology. Microscale data uses molecular-level events, such as protein-small molecule binding and interactions using our protein sciences, biochemical assays, and physics- and chemistry-centric models, to enable insights at the level of target- and protein-interactions and properties.
Macro Scale Data Layers
Macroscale data informs on tissue-, organism-, and population-scale biology, enabling us to connect insights at the molecular (micro) and cellular (meso) levels to the behavior of drug candidates in patients, and to perform “reverse translation” of insights from patient populations to direct the initiation of programs at the beginning of discovery. In 2024, Recursion built investments in macroscale biology in both model organisms for in vivo testing (InVivomics) and human populations, spanning cancer –omic data, population genetics data for non-oncology indications, and real-world clinical data.
30
Table of Contents
InVivomics Data Layer
Recursion’s data layers combine to tell the story that our therapeutics will safely provide benefit to a patient. Currently, in vivo experimentation is necessary to confidently translate the initial insights from high throughput experimentation in biology and chemistry to applications in the real world. Our InVivomics platform removes human toil and bias from animal data collection. Leveraging this platform maximizes data collection while minimizing human effort in key in vivo experimentation areas, in vivo pharmacology and toxicology.
REAL WORLD WORLD MODEL
In 2024, InVivomics produced important data that drove decisions across disease models in fibrosis, neuroscience, and oncology. We integrated tolerability studies into our automated industrial workflows, streamlining the process from hit compound identification to animal model testing through a standardized set of experiments and decision criteria. Our in-house execution of a lung fibrosis model helped accelerate the delivery of a molecule now progressing to clinical trials. We also piloted studies in oncology and neuroscience. In neuroscience, we introduced new endpoint measurements like rotarod and CMAP. Developing this skill set and assessing how digital biomarkers can enhance and accelerate data not only supported an internal project decision but will also play a key role in advancing our partnership projects.
In total in 2024, we ran 62 InVivomic-informed studies at our Milpitas, CA facility. Of those, 27 were mouse tolerability studies, delivering richer data to project teams as they design downstream in vivo pharmacology studies. Across our internal portfolio, 7 projects leveraged this technology to inform dose selection as well as to evaluate impacts on specific tissues of interest. We are also exploring the use of our digital system in rat toxicology studies, evaluating the advantage gained both with the richer constant-monitoring data as well as better connection to our other data layers. We expect a data-driven evaluation in early 2025.
To extract maximum insights from this data, we also built a deep learning model called InVivoPrint V1 (IVP-1) that increases our ability to decode signals coming from these smart cages – detecting liabilities such as inflammation or toxicity earlier than our previous digital biomarker approach. IVP-1 allows us to detect organ toxicities linked to a compound or dose candidate as early as possible during in vivo tolerability studies – and to prioritize new drug candidates for the efficacy phase. Our current InVivomics dataset includes 1 million hours of video; 1 million hours of digital biomarkers such as locomotion, body temperature, wheel speed, and cage humidity levels; 149,000 environment data points, including cage slottings, rack used, rack room, sex, and birth time, as well as a number of other categories.
31
Table of Contents
Figure 20. A Recursion scientist uses our in house dashboard to monitor digital readouts and live video of an ongoing animal experiment.
Human genetics data layer.
At the macroscale, patient omics and observational Real World Evidence (RWE) informs organism-level phenotypes, which is critical for discovering the underlying genetic associations of a particular disease. As an independent data layer, human genetics has proven to be undeniably important for increasing the probability of clinical success. Yet the full value of these patient datasets is limited by the fact that these data are inherently noisy, incomplete, and difficult to collect at a scale that allows for recall of sufficiently rare variants and disease phenotypes. At Recursion, we have the capability to bridge our highly controllable, densely sampled perturbative map data (meso) together with observational patient data (macro) in a joint forward-reverse genetics approach. Using human genetics, we can connect the macroscale down to the mesoscale to inform phenotypic discovery and deliver stronger, more disease-relevant and patient-connected insights. While conversely, by integrating scaled meso data, we in turn increase the power and derive further value out of macro data above standard approaches.
REAL WORLD WORLD MODEL
In 2024, we expanded our real-world macro data layer by partnering with various clinico-genomics companies and acquiring other sources of RWE. In addition to retaining access to over 20PB of de-identified oncology patient data through our partnership with Tempus, we are now partnering with Helix and have scaled access to hundreds of thousands of de-identified non-oncology patient records consisting of longitudinal clinical records paired with Helix’s Exome+® genomic data. We are working with real world data (RWD) providers and continue to augment the foundational macro layer with non-patient trial data such as fit-for-purpose natural history, and both historical and Recursion trial data.
With the integration of these real-world macro and meso data, we built critical components of the Recursion world model to increase the effective power of genetic association tests, inform on patient causality, and enable the precise selection of patient populations based on causal insights. In 2024, we demonstrated that we could extract genetic causal targets from small (24,000) patient data sets that had previously, depending on the target, required patient data sets of hundreds of thousands up to over
32
Table of Contents
one million (3-67x increase in effective power). Recursion believes this to be a new and plausible capital-efficient approach to rare variant discovery.
We also developed a generalizable causal discovery workflow combining macro-meso data features and applied these models for target identification and program initiation purposes. This has led to over 60 genes identified through these causal models that are currently in testing on our validation platform. Beyond discovery, these causal AI models are deployed to identify potential population expansions on current development programs. We have also started testing causal inference models for predicting responsive patient populations to aid in biomarker identification and design optimized patient selection strategies.
Our diversification and expansion of our access to large-scale real-world data, has strengthened our foundation for building world models. This included the integration of electronic health records, claims data, non-patient clinical trial data, and historical trial data. These data sources are being leveraged to advance development through intelligent trial design—optimizing patient selection, trial protocols, and biomarker strategies based on predictive insights; AI-powered clinical trial execution—accelerating patient enrollment via data-driven site selection, automated outreach, and dynamic recruitment optimization; and multi-modal RWE application at scale—combining genomic, clinical, and claims data to inform decision-making across discovery, development, and validation.
Finally, we also accelerated patient enrollment with data-driven site selection and automated site outreach. Looking ahead, we will expand and develop these components, and the industrialized workflows that integrate these into the Recursion OS to drive industrialized clinical development and increase probability of success for our clinical programs.
Meso Scale Data Layers
Mesoscale data at Recursion informs us about cellular biology and serves a unique role: data on cellular systems integrates over biological pathways, revealing information about the multiple effects that individual molecules and targets may mediate to enhance translatability (polypharmacology), while simultaneously offering orders of magnitudes greater sample scale and interventional capability than macro-scale systems. Recursion has built and applied two high-throughput mesoscale assays and data layers, phenomics and transcriptomics.
Phenomics Data Layer
Phenomics measures the morphology of cultured cells grown in laboratory plates. Morphology is a holistic measure of cellular state that integrates changes from underlying layers of cell biology, including gene expression, protein production and modification, and cell signaling, into a single, powerful readout. Image-based -omics can be two to four orders of magnitude more data-dense per dollar than other -omics datasets that focus on these more proximal readouts, enabling us to generate far more data per dollar spent to inform our drug discovery efforts. Phenomics data on genetic and small molecule perturbations forms the backbone of Recursion’s Maps of Biology.
REAL WORLD WORLD MODEL
Our real-world experimental phenomics data has historically been captured using multi-channel fluorescence microscopy; through 2024, we transitioned the platform to acquire live-cell brightfield images, an imaging modality bringing the capability to measure dynamic cellular state across time, rather than at a single timepoint as is typical for fluorescence-based phenomics or sequencing. In 2024, Recursion’s real-world phenomics experimental capabilities scaled to be able to generate up to 13.2 million cell paint images (110 terabytes) or up to 16.2 (135 terabytes) million multi-timepoint brightfield images across up to 2.2 million experiments per week.
Our state-of-the-art machine learning work in phenomics contributes deeply to the Recursion world model. In 2024, Recursion demonstrated the power of scaling laws in machine learning with the training and deployment of Phenom-2, a larger version of the Phenom-1 phenomics model from 2023, making use of the increased computational power of BioHive-2 to improve the detection rate of expressed gene knockouts by 25.7%. We further demonstrated the power of Phenom models on our data by training a Phenom-2-derived model to reconstruct fluorescent Cell Painting images from brightfield data alone, potentially enabling us to directly relate historical Cell Painting data to the brightfield data being collected today and in the future.
33
Table of Contents
Figure 21. Predicting Cell Paint from Brightfield. Phenom-derived models are able to accurately impute fluorescent stains from brightfield-only images. Using paired brightfield and fluorescent phenomics (“Cell Paint”) data, Recursion scientists trained a Phenom-derived model to reconstruct fluorescent images from brightfield data.
Transcriptomics and new –omics data layers
Transcriptomics is a high-dimensional measure of cellular biology distinct from phenomics that assesses gene expression by measuring RNA levels in the cell. Transcriptomics augments our mesoscale data acquisition in three key ways. First, it enables independent replication, at scale, of effects detected in phenomics to verify that they are not morphology-specific artifacts. Second, it offers a route to greater potential interpretability of high-dimensional biological effects by mapping perturbations onto identifiable genetic pathways. Finally, it potentially enables the acquisition of new kinds of biological information, including both effects specific to the transcriptome and new perturbations inaccessible on our phenomics platform.
REAL WORLD WORLD MODEL
We acquire transcriptomics data in the real world using both an internally developed high-density bulk arrayed transcriptomics platform as well as a newly developed pooled single-cell transcriptomics capability. In 2024, we have expanded our arrayed transcriptomics platform capability to enable the sequencing up to 62,000 wells per week and in 2024, generated just under 1 million individual transcriptomes of data. This year, we augmented our platform with the capability to read out pooled perturbations by single-cell RNA sequencing and demonstrated this capability with what we believe to be the world’s first genome-scale CRISPR knockout map in primary human cells. We continue to explore investments in –omics technologies beyond transcriptomics, including but not limited to proteomics and metabolomics.
Transcriptomics represents the first extension beyond phenomics in the Recursion world model. In 2024, we applied Recursion algorithms operating on transcriptomic experiments confirming phenomics to replace time-consuming, disease-specific validation assays with a portfolio-wide multimodal analysis. This analysis demonstrated a 90% ability to predict compounds that failed later disease-relevant assays in internal tests and 60% ability to predict compounds that passed later disease-relevant assays in internal tests. The results of our whole-genome transcriptomic knockout map are now available for internal analysis in the Recursion Data Universe, and in 2025 we anticipate further development of scaled machine learning capabilities on transcriptomic data paralleling our historical development in phenomics.
34
Table of Contents
Figure 22. Phenomics and transcriptomics can provide complementary views of biology. In this figure, relationships between knockouts of genes involved in control of RNA transcription and protein translation are visualized, with relationships from phenomics in the lower triangle and those from our internal genome-scale transcriptomic knockout map in the upper triangle. As a distal readout, phenomics sees similar effects on cellular biology from the loss of either transcription or translation. By contrast, transcriptomics identifies opposite directionality of effect between knockouts of these two classes of genes, potentially offering higher resolution in certain areas of biology.
Micro Scale Data Layers
At the smallest end, microscale data informs on the key chemical and biophysical measurements needed to succeed in drug discovery. This scale covers the molecular-level events, such as the binding events between compounds and their target proteins, as well as the chemical reactions involved in synthesizing and metabolizing these compounds. Three data layers encompass this micro scale. (1) Protein Target, (2) Chemical Data & Automated Synthesis, and (3) ADMET, each encompassing scaled data generation and state-of-the-art AI models to accelerate our design initiatives.
Protein Target Data Layer
Our protein target data layer measures the protein-ligand binding interactions that drive drug discovery. Engaging protein targets with new compounds is a key driver in the development of effective medicines. This data layer encompasses the development of new target-centric functional assays, the automated platform conducting these real-world experiments, as well as our suite of advanced physics-based simulations that yield accurate synthetic data. These insights are captured by our state-of-the-art predictive chemistry models, using this data to guide automated design decisions.
REAL WORLD WORLD MODEL
In 2024, we scaled and automated our experimental bioassay platform, which drives the testing phase of our precision Design-Make-Test-Analyze (DMTA) active learning loop. The platform's key features include assay type diversity, speed of execution, and close integration with software tools to enable autonomous operation. The platform currently supports over 250 diverse biochemical and functional assay types, allowing us to drive diverse, high-quality target-enabled programs. All assay plates are prepared autonomously, without human intervention, with over 50% of assay plate preparations taking place overnight in unattended facilities.
35
Table of Contents
Recursion enhances its automated bioassay platform with absolute and relative binding free energy simulations (ABFE/RBFE) of protein-ligand interactions and protein folding predictions, each serving a different purpose toward enhancing molecule design. ABFE and RBFE are accurate molecular dynamics (MD) simulations of protein-ligand binding interactions, which evaluate new chemical design on structurally enabled targets by computationally determining their affinities to the target. Our RBFE calculations have demonstrated an average accuracy of 1.3 kCal/mol. These evaluations prioritize which designed compounds are made and tested.
From there, protein folding and co-folding predictions are used to locate the ligand binding and mechanistic targets responsible for observed biology activities detected by our meso data layers (phenomics and transcriptomics). In 2024, we connected 1.4 million known active ligands mapped to specific pockets across a synthetic data layer of 3D human protein structures. These relationships are used to identify tentative off-target interactions, binding site (and their key active residues), and to initiate subsequent structure-based modeling, including ABFE and RBFE simulations. Bioactivity assays measured through the bioassay platform and sourced across the Recursion Data Universe are further modelled through a suite of state-of-the-art machine learning models. Notably, activity models are built with our MolGPS foundation model pretrained on thousands of chemical properties and biological activities. Together, each virtual model enhances the “Design” capability of the iterative Design Make Test Learning loop.
Figure 23. Interaction surfaces of a molecule as determined by fragment hotspot analysis. Fragment hotspot maps are used to identify druggable sites on protein surfaces, and to map target similarity across a synthetic data layer of 3D protein structures.
Chemical Data Layer & Automated Synthesis
Our chemical data layer integrates precision design, state-of-the-art molecular property prediction, and fully automated chemical synthesis, with the goal of designing and producing high-quality, differentiated medicines for patients. The precision design element transforms the drug discovery and development process. It replaces the current conventional/conservative approach with an AI-first learning system that is well-suited to the complexities of drug discovery in each step of the process. Our proprietary AI excels at generative molecular design, molecular property prediction and multi-parameter optimization - powering our platform to multiplex design against addressing more complex, desired profiles than conventional approaches, with synthetic accessibility at its core.
REAL WORLD WORLD MODEL
The synthesis-first, generative design approach addresses a significant challenge posed by early generative methods, which frequently yield undesirable molecules that are difficult to synthesize. Our approach represents a substantial advancement in generative design and can be considered a next-generation solution. As we design molecules with AI-driven algorithms, our
36
Table of Contents
platform, informed by our performant property prediction models, guides the design process towards compounds that are not only biologically and physiologically optimized, but also amenable to efficient synthesis on a unique, fully automated platform, suggesting the most cost and time-effective way to make the molecule with sophisticated retrosynthesis that is vendor-logistics-aware.
We capture reaction data, alongside bioassay data, and leverage active learning to ensure that our molecular property and synthesis prediction models continue to learn in step with the evolution of our pipeline, continuously refining our ability to identify the most promising molecules and predict the feasibility and success of future synthetic routes. Recursion’s AI synthesis planning capability shows a 25% improved tractability assessment of AI-generated compounds over competitors and integrates with the Recursion OS. Incorporation of our platform into our discovery processes has resulted in a 35% improvement in design cycle productivity, enabling our design platform to support a broad pipeline.
Figure 24. Recursion's automated chemistry wet lab, a modular system for chemical synthesis preparation, execution, analysis, work-up, and purification.
ADMET Data Layer
A durable truth in drug discovery is the requirement that we have confidence in our prediction of how our drugs will behave in patients. We must enter clinical experimentation assured that we have a reasonable expectation that we will safely deliver benefit to patients. A key part of building that confidence is early testing of a candidate molecule’s pharmacokinetic (PK) properties, which informs the likelihood that the drug will stay in a patient’s body for the right amount of time to be effective. At Recursion we strive to generate critical decision-making data as early as possible. Focusing our testing resources on compounds with higher likelihood to advance accelerates our mission to radically improve lives.
REAL WORLD WORLD MODEL
In 2024, we leveraged our high throughput ADME platform, RADME-01, to generate a bolus of data informing a compound’s likelihood to be viable. The RADME-01 platform semi-autonomously runs a suite of ADME experiments including passive permeability, metabolic stability in liver microsomes, and non-specific protein binding. Our platform has been operational since late 2023 and can evaluate two 384-well plates in all available assays each week. In 2024, we tested 12,209 novel compounds through this system, supporting both Recursion’s internal pipeline and partnered projects. Throughout the year the automation system evolved and refined, allowing for continually less human intervention. A sophisticated in-house built software system intelligently prioritizes compounds and designs experiments to deliver highest value data while avoiding pitfalls related to compound analysis at scale. Drug discovery project teams can use this data to prioritize higher quality compounds for subsequent evaluation, and we’ve also installed pre-set criteria for compound evaluation in our earlier (hit to lead) stage. In this semi-autonomous loop, data is generated and compounds are nominated for in vivo PK testing without human effort or bias. We are evaluating these criteria on a regular basis to ensure we are giving our projects the best chance at advancing high quality projects and compounds.
Another primary use of data generated on our RADME-01 platform is to train predictive models, capable of evaluating designs before synthesizing new molecules. In 2024, we developed an automated machine learning framework that retrains and deploys new RADME-01 assay endpoint models weekly upon availability of new data, as well as several curated ADME property
37
Table of Contents
datasets. These models included a broad set of properties such as intrinsic clearance, non-specific binding, solubility, efflux, drug-drug interactions, and several human PK parameters. Our teams used these models for both internal pipeline and partnership projects both for prioritization of experimentation and prioritization of new molecule synthesis targets. Recently, automated models derived from the RADME model were combined with their counterparts in the Centaur Chemist platform, further enabling rapid design compounds across both internal and partnership projects. Our combined datasets revealed very little direct overlap and have expanded the chemical space our models are trained on.
Figure 25. The Exscientia and Recursion ADME datasets had few overlapping compound matches but were sufficiently representative to train ADME models with equivalent or improved performances across all endpoints.
Scientific agents and industrialized workflows
Agentic systems are artificial intelligence systems, typically large language model-based, that can pursue complex goals with limited human intervention. They make decisions in a context-dependent manner and can learn and adapt through interaction with the world around them. They are perfectly suited to workflow design leveraging the data layer and model inference-based modules of the Recursion OS, automating the design of workflows to meet a human or machine-stated objective.
Our goal at Recursion is to industrialize drug discovery and development through standardization and automation. The early stages of our drug discovery process, filling the T-shaped funnel, benefits from standardization and exploitation of the program-agnostic data universe that Recursion has generated this past decade. Industrializing later stages of drug discovery, where the focus is on molecular design, make and test in a program-specific assay cascade, requires context-dependent automation and exploitation of program-specific data. Agentic systems, with their adaptability and flexibility, provide Recursion with the opportunity to industrialize the drug discovery and development process in its entirety.
The industrialization of drug discovery and development at Recursion is implemented via a series of Industrialized Workflows, each exploiting the Recursion OS to serve both the mission of decoding biology to radically save lives and the pipeline. The first of these workflows, Initiation Workflow, offers a standardized approach to program initiation, the filling of the T-shaped funnel. Its role is to generate and assess disease-gene hypotheses. It succeeds by querying the patient data and maps of biology that reside in the data universe and, with help from a suite of Recursion OS models, hypothesizes which genes are associated with which diseases. Large language model-based approaches are used to annotate these hypotheses with strategic insights collated from external, unstructured sources. Such insights inform the biological relevance, novelty, commercial opportunity, competitor landscape, and the opportunity to differentiate. LLMs are also responsible for assigning a score related to the strength of the hypothesized relationship between each disease-gene pair.
The industrialization of lead optimization requires a different approach. At this stage, a program’s needs can be unique to that program. The assay cascade may be different and the chemical series undergoing optimization will be different. Recursion’s AI precision design platform, Centaur Chemist, enables computational and medicinal chemists to build molecular design workflows that design molecules that are synthesizable on our automated chemical synthesis platform and meet a particular design cycle objective.
38
Table of Contents
Our precision design approach represents a constant interplay of automated data generation from experimentation and intelligent learning systems, embodying virtuous cycles of improvement in chemistry optimization. Through our integrated platform, we can systematically encode the goals and strategy of each molecular design cycle, executing them through a sophisticated combination of scientific technology modules that form cohesive workflows. Compounds progress through synthesis and testing, generating valuable data that are automatically captured in our integrated platform, along with valuable annotations from our medicinal chemists, ready to inform and enhance our predictive modules. This closed-loop system helps translate design objectives into executable workflows, leveraging our 30+ scientific tech modules: from 2D and 3D synthesis-aware generative methods to property prediction models, reinforced with physics-based thinking, and active learning approaches for compound selection. As compounds are generated, synthesized, tested, and evaluated, the platform captures every decision point and experimental outcome, feeding this information back into our models to enhance future design choices. This recursive learning process ensures each iteration becomes more precise and informed than the last, driving the evolution of a more intelligent and efficient drug discovery engine that continuously learns and adapts.
Figure 26. Industrializing lead optimization with an integrated design and synthesis workflow.
Processing and Data Storage Infrastructure
We believe modern drug discovery and development is a data and compute problem – the need to understand pathways, targets, compounds, and mechanisms of actions requires obtaining, synthesizing, or predicting large volumes of data. To store this data in an efficient and low-risk way, Recursion makes use of a combination of cloud storage, and on-premises storage. To process this data efficiently, we bring it close to where the compute will run – either in our HPC datacenter (BioHive) or to our cloud (partnering with Google Cloud). To make this more seamless for our scientists, we have invested in a hybrid storage and compute platform, which enables replication of data and locality of compute to allow us to use these resources as efficiently as possible.
This year we expanded our partnership with Google Cloud to explore generative AI capabilities, including Gemini models, supporting the Recursion OS, driving improved search and access with BigQuery, and helping scale compute resources.
We currently have over tens of petabytes of unique data replicated across our sites for redundancy and resiliency. We use this data to train state of the art (SOTA) foundation models of biology and chemistry and continue to push the limits of what is possible as we invest to scale up beyond Phenom-2 on the biology side, and bring unique Quantum Mechanics (QM) and Molecular Dynamics (MD) data to bear in our chemistry models like MolGPS. We largely train these models using our own supercomputer which consists of two generations of DGX SuperPod, with 504 H100s, and 320 A100s, and over 65 TB of VRAM.
39
Table of Contents
Figure 27. BioHive-2 is Recursion’s new NVIDIA DGX SuperPOD AI supercomputer, powered by 63 DGX H100 systems with a total of 504 NVIDIA H100 Tensor Core GPUs interconnected by NVIDIA Quantum-2 InfiniBand networking. This NVIDIA-powered AI supercomputer results in over four times faster speeds than Recursion’s original supercomputer, BioHive-1, in benchmark performance tests. Based on available data, BioHive-2 is the fastest supercomputer wholly owned and operated by any pharmaceutical company worldwide.
Bringing it together – Combining data scales to drive value
The vision of the Recursion OS is to integrate data and insights across biological scales to build a comprehensive and predictive World Model that deciphers biology with unprecedented efficiency. Our ability to generate and leverage real-world data at scale—across macro, meso, and micro levels—has already begun to reshape how we approach drug discovery. By combining these layers through the Recursion OS, we are moving toward an industrialized, AI-driven system that not only accelerates therapeutic discovery, but we also believe will help us to increase the probability of success in clinical development.
Our first demonstrations of this approach have successfully integrated macroscale patient data with mesoscale phenomics, allowing us to extract genetic causal targets from datasets previously considered too small for statistical power. This methodology enables a capital-efficient approach to rare variant discovery, increasing our ability to identify novel drug targets that are deeply connected to human disease, and it is just the beginning.
At the mesoscale, we have advanced our phenomics and transcriptomics capabilities to provide high-throughput, high-dimensional insights into cellular biology. By connecting these layers with microscale molecular interactions, such as protein-ligand binding and ADMET properties, we enhance the interpretability and mechanistic understanding of our drug candidates. We believe that bridging across all three scales—macro, meso, and micro—will unlock a deeper understanding of human biology and significantly improve the efficiency of drug discovery and development.
One of the most transformative outcomes of our integrated approach will be the ability to construct a Virtual Cell—an AI-powered system that simulates biological responses at scale. Traditionally, drug discovery has been an experiment-driven process where models are built from data collected in the lab. At Recursion, we are reversing this paradigm: our World Model is now driving the generation of new hypotheses, with real-world experimentation serving to validate the most promising insights. By iteratively refining these AI-driven predictions with physical experiments, we are creating a feedback loop that accelerates learning and reduces reliance on trial-and-error experimentation.
As this approach matures, we envision a future where the Virtual Cell serves as a comprehensive digital twin for human biology, enabling us to model drug interactions, disease progressions, and therapeutic interventions in silico before ever entering the lab. This shift—from experiment-first to model-first—has the potential to revolutionize how drugs are discovered, reducing both time and cost while significantly improving success rates.
Scaling Beyond What Was Previously Possible in 2024
In 2024, Recursion expanded its ability to industrialize drug discovery by enhancing the OS’s automation, scalability, and machine learning capabilities. Key milestones included:
•The launch of BioHive-2, the most powerful supercomputer owned by any biopharma company, enabling the training of industry-leading foundation models like Phenom-2, MolPhenix, and MolGPS.
•The integration of Exscientia’s automated chemistry platform, which has already generated over 500 custom molecules in under nine days per cycle.
•The successful completion of the world’s first whole genome neuronal phenotype map (Neuromap) in partnership with Roche and Genentech, representing a significant leap forward in neuroscience drug discovery.
40
Table of Contents
•The augmentation of our real-world data layer with hundreds of thousands of patient records through new partnerships with Helix and Tempus, dramatically improving our ability to connect patient-level insights with early-stage discovery.
•The rapid expansion of our transcriptomics capabilities, surpassing 1 million whole transcriptomes sequenced in a single year, reinforcing our ability to generate multimodal insights.
•The development of InVivoPrint V1 (IVP-1), an advanced deep learning model that enhances our ability to detect organ toxicities and prioritize drug candidates with greater precision.
The Road Ahead
Recursion is at the forefront of a new era in drug discovery, where data, AI, and automation converge to redefine the boundaries of what is possible. The continued evolution of the Recursion OS will focus on:
•Further expansion of multimodal AI models that integrate patient-level insights with cellular and molecular data to refine drug target selection.
•Greater automation in preclinical validation through advances in high-throughput biology and AI-driven chemistry.
•Industrialized clinical development leveraging AI-powered trial design and patient selection to increase the probability of success.
•Scaling our Virtual Cell approach to predict and validate therapeutic interventions with unparalleled accuracy.
As we move forward, we remain committed to the mission that has guided Recursion from the beginning: to decode biology to radically improve lives. With our unique combination of scaled experimentation, AI-driven insights, and industry-leading automation, we are not just advancing drug discovery—we are fundamentally redefining its future.
Our Pipeline
Programs in our internal pipeline are built on unique biological and chemical insights surfaced through the Recursion OS where:
I.The etiology of the disease is well defined, but the subsequent impacts of the disease are generally obscure and/or the primary targets are typically considered undruggable.
II.There is a high unmet medical need, no approved therapies, or significant limitations to existing treatments.
Following the combination with Exscientia, we have expanded our internal and partnered portfolio, adding multiple programs across oncology, immunology, rare diseases, neuroscience, and more. Beyond our 10 clinical and preclinical programs, we are advancing 10+ next-gen discovery programs for further development.
Figure 28. The power of our Recursion OS exemplified by our expansive therapeutic pipeline. 1Includes non-small cell lung cancer (NSCLC), colorectal cancer, breast cancer, pancreatic cancer, ovarian cancer, head and neck cancer; 2Joint venture with Rallybio.
41
Table of Contents
Clinical Programs in Oncology
REC-617 – Advanced Solid Tumors
REC-617 is a potential best-in-class, potent and selective oral small molecule inhibitor of CDK7 with demonstrated activity in preclinical studies. CDK7 controls cell cycle progression and gene transcription, often overexpressed in advanced stage cancers reliant on transcriptional pathways. This program utilized our generative AI and active learning platform to optimize molecule design, including non-covalent binding and ADME/PK for rapid absorption. This rapid design cycle enabled us to synthesize 136 novel compounds and select REC-617 as our lead candidate in under 11 months.
A multicenter, open-label, Phase 1/2 (ELUCIDATE) monotherapy dose escalation (QD and BID) study is currently ongoing in advanced solid tumors. In December 2024, results from the initial 19 patients (18 response-evaluable at the time of cutoff) were presented at an AACR Special Conference in Cancer Research. REC-617 monotherapy demonstrated signs of preliminary efficacy. One heavily pre-treated ovarian cancer patient achieved a confirmed durable partial response (PR), which correlated with significant reductions in clinical tumor markers (CA125 and TK1). Four additional patients achieved durable stable disease (SD) as their best response. REC-617 was generally well-tolerated, with adverse events predominantly low grade, on-target, and reversible upon treatment cessation. The MTD was not reached and there were no treatment-related discontinuations.
Monotherapy dose escalation (QD and BID) remains ongoing, and we expect to initiate combination studies in the first half of 2025. We also expect to provide additional data updates from the Phase 1 in 2025.
REC-1245 – Biomarker-enriched Solid Tumors and Lymphoma
REC-1245 is a first-in-class, novel, potent, and selective molecular glue degrader of RBM39, a critical RNA-binding protein involved in alternative splicing and DNA damage repair (DDR) pathways. Leveraging the Recursion OS, we discovered that genetic knockout of RBM39 can phenotypically mimic CDK12 loss – a validated DDR target –without impacting CDK13. To our knowledge, we were the first to report this novel biological insight. Utilizing our phenomics based platform for SAR, we synthesized 204 candidates and advanced this program from target ID to IND-enabling studies in 18 months (vs. industry average of 42 months).
Preclinical data confirmed strong anti-tumor activity, including tumor regressions in a BRCA-proficient ovarian cancer model, minimal off-target effects, and no CDK12 kinase inhibition. With over 100,000 addressable patients in the US and EU5 each year, REC-1245 has the potential to be a novel therapy in a biomarker-enriched advanced solid tumors and lymphoma patient population – either as a monotherapy and/or in combination regimens.
Following IND clearance in September 2024, we initiated a Phase 1/2 (DAHLIA) study in December 2024 to evaluate the safety, tolerability, PK/PD, and preliminary efficacy of REC-1245 in unresectable, locally advanced, or metastatic cancers. The trial is currently enrolling at three US sites and includes a biomarker-enriched population that may benefit most from targeted RBM39 degradation. We expect to share an update on the Phase 1 dose-escalation portion of the study in the first half of 2026.
REC-3565 – Relapsed / Refractory B-cell Malignancies
We are advancing REC-3565, our reversible allosteric potential best-in-class MALT1 inhibitor, for the treatment of patients with relapsed or refractory B-cell malignancies. A variety of mutations seen in lymphomas induce constitutive MALT1 protease activation, leading to aberrant NF-κB signaling that drives survival and proliferation of B-cell tumors. Key preclinical data demonstrates sustained anti-tumor activity as a single-agent or in combination with BTK inhibitors.
We leveraged physics-based predictive modelling using our molecular dynamics toolkit and AI-powered hotspot analysis to deliver a candidate with lower predicted safety risk in the clinic. We synthesized 344 novel compounds and advanced this program from hit ID to lead candidate in 15 months.
The molecule’s unique profile minimizes UGT1A1 inhibition risk, demonstrating superior target selectivity compared to oral competitors, both of which reported treatment-related hyperbilirubinemia in early Phase 1/2 studies. As a result, REC-3565’s enhanced selectivity supports the potential for a more favorable therapeutic index not only as a monotherapy, but also in combinations with BTK and BCL2 inhibitors. A multicenter, open-label, dose escalation Phase 1 study (EXCELERIZE) cleared a CTA by the MHRA in December 2024. We expect to dose the first patient in the first half of 2025.
REC-4539 – Small Cell Lung Cancer
REC-4539 is reversible CNS penetrant, orally bioavailable, and potential best-in-class inhibitor of LSD1. LSD1 is an epigenetic enzyme that removes methyl groups from histones to control gene expression. SCLC is particularly dependent on LSD1 to maintain a neuroendocrine phenotype that drives tumor cell survival in this aggressive lung cancer subtype. Preclinical studies demonstrate that REC-4539 shows anti-tumor activity in SCLC human xenografts with limited impact on platelets.
Our program used multi-parameter optimization to design a unique candidate combining reversibility with CNS penetration. We synthesized 414 novel candidates to arrive at our lead candidate in 22 months. Following IND clearance in January 2025, we
42
Table of Contents
expect to initiate a multicenter, open-label Phase 1/2 trial (ENLYGHT) in the first half of 2025. We plan to target an SCLC patient population as well as additional biomarker-selected cancers following the dose escalation portion.
Clinical Programs in Rare Diseases
REC-994 – Cerebral Cavernous Malformation
We are developing REC-994, an orally bioavailable small molecule superoxide scavenger, as a first-in-disease opportunity for symptomatic cerebral cavernous malformations (CCM). CCMs are rare vascular anomalies marked by abnormal capillary-venous structures, recurrent lesions, and stroke-like symptoms. REC-994 was discovered using the earliest version of Recursion’s comprehensive drug discovery platform. In an unbiased CCM2 loss of function phenotypic screen, REC-994 demonstrated concentration dependent rescue and was advanced into preclinical studies. In animal models of CCM, REC-994 reduced the burden of CCM lesions by ~50%. In addition, REC-994 also reduced the vascular permeability defects in CCM2-deficient mice, which is critical in CCM pathology. This data supported the clinical development of REC-994, the first industry-sponsored trial for CCM.
In late 2020, initial results from the clinical program were reported and a randomized Phase 2 trial of REC-994 (SYCAMORE) was initiated in March 2022 in patients with symptomatic CCM. In April 2024, the Phase 1 SAD/MAD study was published. Initial results in September 2024 showed REC-994 met its primary endpoint of safety with encouraging trends in preliminary efficacy. The drug was well-tolerated, with no treatment-related discontinuations or Grade 3-4 adverse events reported. We presented the Phase 2 study data as a late-breaking oral presentation at the International Stroke Conference, or ISC, annual meeting in February 2025. The data presented showed signs of safety and efficacy as follows:
•REC-994 met the primary endpoint of safety and tolerability in CCM patients with no treatment-related discontinuations, SAEs or Grade ≥ 3 adverse events related to study drug
•No new safety signals observed, with the incidence of adverse events comparable across arms
•No treatment-related adverse events that led to discontinuations
•50% of patients on REC-994 400 mg (n=20) achieved a reduction in total lesion volume versus 28% of patients in placebo (n=18) and 24% of patients on REC-994 200 mg (n=17)
•Trends towards improvement and/or stabilization of symptoms for patients treated with REC-994 400 mg (n=19) compared to placebo, which observed trends towards functional decline, based on changes in the Modified Rankin Scale (mRS) score from baseline to 12 months
Similar trends of exploratory efficacy (lesion volume reduction and functional outcome improvement) were seen in the cohort of patients with brainstem lesions treated with REC-994 400 mg
Most (80%) patients who completed at least 12 months of treatment in the Phase 2 study elected to continue into the long-term extension (LTE) portion of the trial. As of December 31, 2024, the LTE portion is ongoing. As there are no therapeutic options for patients with symptomatic CCM, we plan to seek regulatory guidance from the FDA and additional health authorities on a path forward for this potential first-in-disease program. We expect to share updates on next steps in 2025.
REC-4881 – Familial Adenomatous Polyposis
We are developing REC-4881, a highly potent and selective, potential best-in-class MEK1/2 inhibitor, for familial adenomatous polyposis (FAP). FAP is a genetic condition characterized by the development of adenomas throughout the GI tract. It is an orphan disease caused by inactivating mutations in APC, with most patients undergoing prophylactic colectomy due to nearly 100% likelihood of CRC by age 60.
During a collaboration with Takeda, we leveraged machine vision and automated analysis to quantify hundreds of cellular parameters linked to APC siRNA knockdown. We screened numerous compounds in this genetic background for 24 hours and identified REC-4881 as a potent molecule that rescued the phenotype in a concentration dependent manner. In preclinical studies, REC-4881 demonstrated over 1,000-fold selectivity in APC-mutant tumor cell lines and effectively inhibited spheroid growth and organization. In the APCmin mouse model of FAP, REC-4881 showed up to a 70% reduction in total polyps, surpassing celecoxib’s 30% reduction, highlighting its potential as a highly selective and efficacious therapy for FAP.
In April 2022, the IND was reactivated and in September 2022, the Phase 1b/2 trial (TUPELO) of REC-4881 was initiated. As of December 31, 2024, Part 1 of the study is complete, and Part 2 remains ongoing. We expect to share safety and preliminary efficacy data in the first half of 2025.
REC-2282 – Neurofibromatosis Type 2
We are developing REC-2282, a CNS penetrant, potential best-in-class pan-HDAC inhibitor, for neurofibromatosis type 2 (NF2). NF2 is a rare genetic disease caused by loss of function mutations in the NF2 gene which leads to deficiencies in the tumor suppressor protein merlin.
43
Table of Contents
REC-2282 was identified as a potential therapeutic capable of rescuing HUVEC cells treated with NF2 siRNA and subsequently in-licensed from Ohio State Innovation Foundation in December 2018. We initiated the POPLAR study, an adaptive, randomized, multicenter Phase 2/3 trial in June 2022, with the first patient dosed in October 2022. In November 2024, we announced that the trial was fully enrolled in the Phase 2 portion.
As of December 31, 2024, Phase 2 data is maturing, and we expect to share the results of the futility analysis (PFS6 rate) in the first half of 2025.
REC-3964 – Prevention of Recurrent C. difficile infection
We are developing REC-3964, a non-microbial, orally bioavailable, potential first-in-class C. difficile (C. diff) toxin B selective inhibitor for the prevention of recurrent Clostridioides difficile infection (rCDI). C. diff toxin B disrupts the tight junctions in colonic cells and increases vascular permeability, leading to a leaky gut. REC-3964 is Recursion’s first new chemical entity to reach the clinic and binds and blocks the catalytic activity of the toxin's innate glucosyltransferase, while sparing the host. In a human disease relevant C. diff. hamster model, REC-3964 demonstrated a significant difference in the probability of survival versus bezlotoxumab alone.
Our program leveraged an ML-aided conditional phenotypic drug screen in human cells and identified novel mechanisms that mitigated the effect of C. diff. toxin B treatment. Through orthogonal validation screens, precursors to REC-3964 emerged as promising substrates for further advancement.
In June 2024, we presented Phase 1 data in healthy volunteers at the 6th Edition of World Congress on Infectious Diseases in Paris. In October 2024, we initiated a Phase 2 open-label, randomized, 3-arm study (ALDER) to evaluate the rate of recurrence in patients with a high-risk of CDI, who have achieved symptom resolution following treatment with oral vancomycin for 14 days. We expect to share initial results from the Phase 2 study in the first quarter of 2026.
Details on preclinical programs (e.g., ENPP1 inhibitor, Target Epsilon) will be shared in the next section.
44
Table of Contents
Deep Dive into Clinical and Select Preclinical Programs
REC-617 for Advanced Solid Tumors – Phase 1/2
REC-617 is an orally bioavailable, cyclin-dependent kinase 7 (CDK7) inhibitor currently under development for the treatment of advanced solid tumors. Inhibiting CDK7 targets both cell cycle dysregulation and transcriptional "addiction", which are hallmarks of multiple aggressive cancers including, but not limited to, CDK4/6 resistant breast cancer, ovarian cancer, and other solid tumors. There are currently no CDK7 inhibitors approved by the FDA. ELUCIDATE, a Phase 1/2 open-label, multicenter, safety, PK, PD and preliminary efficacy study is currently underway. Interim Phase 1 safety, PK, PD, and efficacy data were shared in the fourth quarter of 2024. We expect to initiate combination studies in the first half of 2025.
Disease Overview
The importance of cell cycle inhibitors in oncology has been established with CDK4/6 inhibitors, which generated approximately $10.5 billion in sales in 2023. Aberrant CDK7 overexpression is common in many cancer indications and associated with poor prognosis. CDK7 presents an opportunity to improve treatment outcomes over CDK4/6 inhibitors due to CDK7’s dual role in cell cycle and transcription. Potential specific indications include non-small cell lung cancer (NSCLC), colorectal cancer, breast cancer, pancreatic cancer, ovarian cancer, head and neck cancer for which we estimate an addressable population of approximately 185,000 drug-treatable patients per year in the US and EU5.
Insight from Recursion OS
CDK7 inhibitor development has faced significant challenges, primarily due to off-target effects and suboptimal pharmacokinetics. Previous attempts often employed covalent binding mechanisms or exhibited poor oral bioavailability, leading to undesirable side effects in the clinic. Current candidates in development for CDK7 feature covalent binding or extended half-lives potentially resulting in substantial on-target toxicity. In addition, the reversible inhibitors under investigation are transporter substrates, likely compromising their absorption and exacerbating gastrointestinal adverse events. These limitations underscore the critical need for novel CDK7 inhibitor designs that optimize both safety and efficacy profiles.
Leveraging our AI-driven multi-parameter optimization approach, we identified critical design limitations in existing CDK7 inhibitors. This insight led to an improved target product profile and a novel molecule design. REC-617 is an orally bioavailable, potent and selective CDK7 inhibitor with enhanced oral bioavailability. It has a non-covalent, reversible mechanism of action, and a predicted shorter human half-life compared to other drugs in development. These characteristics potentially offer an improved therapeutic index, less off-target effects, and more consistent absorption.
Preclinical
REC-617 has demonstrated strong anti-tumor activities in preclinical studies and in vivo experiments showed potent tumor regression across multiple solid tumor types. Notably, in the OVCAR3 ovarian cancer xenograft model as shown below, complete tumor regression was observed in all 8 mice treated with 10 mg/kg by Day 27. Importantly, no significant body weight loss was observed across treatment arms. Mouse PK studies revealed that maintaining 8-10 hours of CDK7 IC80 coverage resulted in potent tumor regression with minimal side effects, while coverage beyond 10 hours led to significant body weight loss. This defined an optimal therapeutic window that guided target efficacious exposures in the clinic.
45
Table of Contents
Figure 29. REC-1245 anti-tumor activity and PK in preclinical tumor models. (Left) REC-617 induces tumor regression in the OVCAR3 cell line derived xenograft mouse model. N=8, 28 days of treatment, REC-617 administered QD PO. (Right) REC-617 administration results in 8-10 hours of therapeutic coverage at IC80. PK studies conducted in CD1 mice, single-dose administration. >10 hr IC80 results in significant body weight loss.14,15
Clinical
In the third quarter of 2023, we initiated a Phase 1/2 open-label, multicenter study (ELUCIDATE) in patients with advanced solid tumors, with the design shown in the figure below. Currently, monotherapy dose escalation (QD and BID) is ongoing, with combination study initiation expected in the first half of 2025.
Figure 30. ELUDICATE study design. Phase 1/2 trial design to assess the safety, PK, exploratory PD, and efficacy of REC-617 in patients with advanced solid tumors.
14 Besnard, et al. (2022). AI-driven discovery and profiling of GTAEXS-617, a selective and highly potent inhibitor of CDK7 [abstract]. AACR; Cancer Res 2022;82(12_Supplement): 3930.
15 Hallett, et al. (2024). Overcoming traditional design limitations with AI-based discovery. AACR Special Conference in Cancer Research: Optimizing Therapeutic Efficacy and Tolerability through Cancer Chemistry; Plenary Session 1
46
Table of Contents
In December 2024, we presented results from the initial 18 response evaluable patients at an AACR Special Conference in Cancer Research. REC-617 was well-tolerated with predominantly Grade 1-2 adverse events, no treatment-related discontinuations, and fewer GI side-effects than reported for other CDK7 inhibitors. Dose escalation (QD and BID) is ongoing, and the maximum tolerated dose (MTD) has not been reached. PK was dose linear and exceeded the CDK7 IC80 with rapid absorption (Tmax 0.5–2h) and short t1⁄2 (5–6h). Robust target engagement was also observed with rapid increases in POLR2A (3-4x), which normalized within 24 hours. These data are shown in the figure below.
Figure 31. REC-617 clinical plasma pharmacokinetics and pharmacodynamics. (Left) REC-617 plasma concentration, in a dose-linear fashion, well above CDK7 IC80 at peak and well below CDK2 IC80, indicating a broad therapeutic window for the study drug. (Right) POL2RA expression data (PD), which is associated with tumor regression, shows a rapid pharmacodynamic effect and short half-life.16,17
Encouraging antitumor activity included a confirmed partial response (PR), in a heavily pre-treated metastatic ovarian cancer patient, with a durable response that was maintained for more than 6 months of treatment. LDH levels were also normalized, and reductions were observed in CA125 (-44%) and TK1 (-68%). Four additional patients achieved the best response of stable disease (SD) lasting up to six months.
Competitors
We are aware of five active CDK7 inhibitor programs in clinical development:
•Samuraciclib (Carrick Therapeutics): In Phase 2 as a monotherapy and in a range of combination studies
•SY-5609 (Syros Pharmaceuticals): Completed Phase 1 monotherapy and in Phase 1/1b in combination with atezolizumab
•Q-901 (Qurient): In Phase 1/2 in monotherapy and combination with PD-1 inhibitors in solid tumors
•TY-2699a (TYK Medicines): In Phase 1 trial in China only
•EOC-237 (EOC Pharma): In Phase 1 trial in China only
REC-1245 for Solid Tumors and Lymphoma – Phase 1/2
REC-1245 is a novel, potent and selective molecular glue degrader of RNA-binding motif protein 39 (RBM39) currently under development for the treatment of biomarker-enriched solid tumors and lymphoma. There are currently no RBM39 degraders approved by the FDA. Following IND clearance in September 2024, we initiated a Phase 1/2 open-label, multicenter study (DAHLIA) to evaluate the safety, tolerability, PK, PD, RP2D, and preliminary efficacy of REC-1245. With the first patient dosed in December 2024, we expect to share an update on the program in the first half of 2026.
16 Papadopoulos, et al. (2020). EORTC-NCI-AACR (ENA) Symposium
17 Hallett, et al. (2024). Overcoming traditional design limitations with AI-based discovery. AACR Special Conference in Cancer Research: Optimizing Therapeutic Efficacy and Tolerability through Cancer Chemistry; Plenary Session 1
47
Table of Contents
Disease Overview
Alternative splicing and RNA-binding proteins (RBPs) have recently emerged as attractive therapeutic targets for cancer due to their critical roles in the regulation of post-transcriptional modifications, impacts on DNA damage repair pathways, and modulation of cell cycle functions. Recent studies have revealed that RBM39 is an unexpected target of aryl sulfonamides, which can function as molecular glue degraders by forming a ternary complex with RBM39 and the E3 ubiquitin ligase receptor DDB1 and CUL4 associated factor 15 (DCAF15). Additionally, clinical trials have shown that aryl sulfonamides were well tolerated with modest anti-tumor activity seen across a variety of cancers. These findings suggest that RBM39 degraders may show promise as targeted cancer therapies, but the lack of predictive biomarkers and an inadequate understanding of RBM39 biology has limited their therapeutic potential. With over 100,000 addressable patients, with biomarker-enriched solid tumors and other select histologies in the US and EU5 each year, REC-1245 has the potential to be used as a single agent or in combination with chemotherapy and/or immunotherapy.
Insight from Recursion OS
Reports suggest that genetic or pharmacologic depletion of CDK12 can reduce the expression of several genes involved in the homologous recombination repair pathway such as BRCA1 and BRCA2, inducing a BRCA-like phenotype and DDR response. Thus, CDK12 has received considerable interest as a therapeutic target and tumor biomarker for HR-proficient cancers. Despite reports of functional redundancy, we observed that the genetic knockout of CDK12 could be clearly distinguished phenotypically from that of CDK13. Using map-based inference to characterize and relate cellular phenotypes, we identified RBM39 as an alternative target that selectively mimics CDK12 loss, but not CDK13, providing a novel approach for targeting CDK12 biology while circumventing any toxicities that may arise due to CDK13. We subsequently discovered REC-1245 as an RBM39 molecular glue degrader that closely mimics the phenotypic loss of CDK12 and RBM39, but not CDK13. Functionally, REC-1245 treatment globally impacts the expression of many DDR genes but does so in a CDK12 independent manner.
Figure 32. Inferred map relationships between CDK12, CDK13, RBM39 and REC-1245. Map representation demonstrates a high degree of phenotypic similarity between CDK12, RBM39, and multiple concentrations of REC-1245. CDK13 shows little or no functional similarity to CDK12, RBM39, or any concentration of REC-1245.
Preclinical
REC-1245 is a potent, potential first-in-class RBM39 molecular glue degrader with compelling preclinical activity. It showed no significant in vitro safety concerns (CEREP, hERG), no CDK12 kinase activity, and minimal ITGA2 liability – an off-target effect seen with prior RBM39 degraders. As shown in the figures below, REC-1245 demonstrated strong antitumor activities as a single-agent, including tumor regression in an ovarian cancer BRCA-proficient, p53 mutant, OVK18 in vivo cell line derived xenograft (CDX) model. In addition, dose-dependent anti-tumor activity correlated with increases in RBM39 degradation confirming target engagement and an exposure-response-efficacy relationship.
48
Table of Contents
Figure 33. REC-1245 single-agent activity and target engagement. REC-1245 single-agent activity and target engagement. (Left) REC-1245 administered BID PO at doses noted for 15 days. N=8 mice per group. (Right) Percent RBM39 degradation (PD) evaluated at REC-1245 doses noted after 5 days BID oral administration of REC-1245. N=3 mice per group.18