rxrx-20211231
Securities registered pursuant to Section 12(b) of the Act:
Title of each class Trading symbol(s) Name of each exchange on which registered
Class A Common Stock, par value $0.00001 RXRX Nasdaq Global Select Market
Large accelerated filer ☐ Non-accelerated filer x
Accelerated filer ☐ Smaller reporting company ☐
Emerging growth company x
Recursion OS
OncologyRare DiseaseNeuroscienceFibrosisInflammation and Immunology
Table of Contents
TABLE OF CONTENTS
Page
Part I
Item 1. Business 10
Item 1A. Risk Factors 95
Item 1B. Unresolved Staff Comments 139
Item 2. Properties 139
Item 3. Legal Proceedings 139
Item 4. Mine Safety Disclosures 139
Part II
Item 6. [Reserved] 142
Item 7A. Quantitative and Qualitative Disclosures About Market Risk 153
Item 8. Financial Statements and Supplementary Data 155
Item 9A. Controls and Procedures 182
Item 9B. Other Information 182
Item 9C. Disclosure Regarding Foreign Jurisdictions that Prevent Inspections 182
Part III
Item 10. Directors, Executive Officers and Corporate Governance 183
Item 11. Executive Compensation 183
Item 14. Principal Accounting Fees and Services 183
Part IV
Item 15. Exhibits and Financial Statement Schedules 184
3
Table of Contents
PART I
RISK FACTOR SUMMARY
Below is a summary of the principal factors that make an investment in the common stock of Recursion Pharmaceuticals, Inc. (Recursion, the Company, we, us, or our) risky or speculative. This summary does not address all of the risks we face. Additional discussion of the risks summarized below, and other risks that we face, can be found in the section titled “Item 1A. Risk Factors” in this Annual Report on Form 10-K.
•We are a clinical-stage biotechnology company with a limited operating history. We have no products approved for commercial sale and have not generated any revenue from product sales.
•Our drug candidates are in preclinical or clinical development, which are lengthy and expensive processes with uncertain outcomes and the potential for substantial delays.
•We have incurred significant operating losses since our inception, we expect to incur substantial and increasing operating losses for the foreseeable future, and we may not be able to achieve or maintain profitability.
•Our mission is broad and expensive to achieve and we will need to raise substantial additional funding, which may not be available on commercially reasonable terms or at all.
•We expect to finance our cash needs for the foreseeable future potentially through a combination of private and public equity offerings and debt financings, as well as strategic collaborations. If we are unable to raise capital when needed, we would be forced to delay, reduce, or eliminate at least some of our product development programs and other activities, and to possibly cease operations.
•Raising additional capital entails risks, including that it may adversely affect the rights, or dilute the holdings, of our existing stockholders; increase our fixed payment obligations; require us to relinquish rights to our technologies or drug candidates; and/or divert management’s attention from our core business.
•If we are unable to establish additional strategic collaborations on commercially reasonable terms or at all, or if current or future collaborations are not successful, we may have to alter our drug development plans.
•We or our current and future collaborators may never successfully develop and commercialize drug candidates, or the market for approved drug candidates may be less than anticipated, which in either case would materially and adversely affect our financial results and our ability to continue our business operations.
•Our approach to drug discovery is unique and may not lead to successful drug products for various reasons, including potential challenges identifying mechanisms of action for our candidates.
•Although we intend to explore other therapeutic opportunities in addition to the drug candidates we are currently developing, we may fail to identify viable new candidates or we may need to prioritize candidates and, as a result, we may fail to capitalize on profitable market opportunities.
•We may experience delays in initiating and completing clinical trials, including due to difficulties in enrolling patients or maintaining compliance with trial protocols, or our trials may produce inconclusive or negative results.
•If we are unable to obtain or there are delays in obtaining regulatory approvals for our drug candidates in the U.S. or other jurisdictions, or if approval is subject to limitations, we will be unable to commercialize, or delayed or limited in commercializing, the products in that jurisdiction and our ability to generate revenue may be materially impaired.
•Our quarterly and annual operating results may fluctuate significantly due to a variety of factors, a number of which are outside our control or may be difficult to predict, which could cause our stock price to fluctuate or decline.
4
Table of Contents
•If we are not able to develop new solutions and enhancements to our drug discovery platform that keep pace with technological developments, or if we experience breaches or malfunctions affecting our platform, our ability to identify and validate viable drug candidates would be adversely impacted.
•Third parties that provide supplies or equipment, or that manufacture our drug products or drug substances, may not provide sufficient quantities at an acceptable cost or may otherwise fail to perform.
•We or third parties on which we depend may experience system failures, cyber-attacks, and other disruptions to information technology or cloud-based infrastructure, which could harm our business and subject us to liability for disclosure of confidential information.
•Force majeure events, such as the COVID-19 pandemic, a natural disaster, global political instability, or warfare, could materially disrupt our business and the development of our drug candidates.
•If we are unable to adequately protect and enforce our intellectual property rights, including obtaining and maintaining patent protection for our key technology and products that is sufficiently broad, our competitors could develop and commercialize technology and products similar or identical to ours and our ability to successfully commercialize our technology and products may be impaired.
•If we are unable to protect the confidentiality of our trade secrets and know-how, our business and competitive position may be harmed.
•If we fail to comply with our obligations in the agreements under which we collaborate with and/or license intellectual property rights from third parties, or otherwise experience disruptions to our business relationships with our partners, we could lose rights that are important to our business.
•We face substantial competition, which may result in others discovering, developing, or commercializing competing products before we do.
•If we are unable to attract and retain key executives, experienced scientists, and other qualified personnel, our ability to discover and develop drug candidates and pursue our growth strategy could be impaired.
•We are subject to comprehensive statutory and regulatory requirements, noncompliance with which may delay or prevent our ability to market our products or result in fines or other liabilities.
Cautionary Note Regarding Forward-Looking Statements
This Annual Report on Form 10-K contains “forward-looking statements” about us and our industry within the meaning of Section 27A of the Securities Act of 1933, as amended, and Section 21E of the Securities Exchange Act of 1934, as amended. All statements other than statements of historical facts are forward-looking statements. In some cases, you can identify forward-looking statements by terms such as “may,” “will,” “should,” “would,” “expect,” “plan,” “anticipate,” “could,” “intend,” “target,” “project,” “contemplate,” “believe,” “estimate,” “predict,” “potential,” or “continue” or the negative of these terms or other similar expressions. Forward-looking statements contained in this report may include without limitation those regarding:
•our research and development programs
•the initiation, timing, progress, results, and cost of our current and future preclinical and clinical studies, including statements regarding the design of, and the timing of initiation and completion of, studies and related preparatory work, as well as the period during which the results of the studies will become available;
•the ability of our clinical trials to demonstrate the safety and efficacy of our drug candidates, and other positive results;
•the ability and willingness of our collaborators to continue research and development activities relating to our development candidates and investigational medicines;
5
Table of Contents
•future agreements with third parties in connection with the commercialization of our investigational medicines and any other approved product;
•the timing, scope, and likelihood of regulatory filings and approvals, including the timing of Investigational New Drug applications and final approval by the U.S. Food and Drug Administration, or FDA, of our current drug candidates and any other future drug candidates, as well as our ability to maintain any such approvals;
•the timing, scope, or likelihood of foreign regulatory filings and approvals, including our ability to maintain any such approvals;
•the size of the potential market opportunity for our drug candidates, including our estimates of the number of patients who suffer from the diseases we are targeting;
•our ability to identify viable new drug candidates for clinical development and the rate at which we expect to identify such candidates, whether through an inferential approach or otherwise;
•our expectation that the assets that will drive the most value for us are those that we will identify in the future using our datasets and tools;
•our ability to develop and advance our current drug candidates and programs into, and successfully complete, clinical studies;
•our ability to reduce the time or cost or increase the likelihood of success of our research and development relative to the traditional drug discovery paradigm;
•our ability to improve, and the rate of improvement in, our infrastructure, datasets, biology, technology tools, and drug discovery platform, and our ability to realize benefits from such improvements;
•our expectations related to the performance and benefits of our BioHive-1 supercomputer;
•our ability to realize a return on our investment of resources and cash in our drug discovery collaborations;
•our ability to scale like a technology company and to add more programs to our pipeline each year;
•our ability to successfully compete in a highly competitive market;
•our manufacturing, commercialization, and marketing capabilities and strategies;
•our plans relating to commercializing our drug candidates, if approved, including the geographic areas of focus and sales strategy;
•our expectations regarding the approval and use of our drug candidates in combination with other drugs;
•the rate and degree of market acceptance and clinical utility of our current drug candidates, if approved, and other drug candidates we may develop;
•our competitive position and the success of competing approaches that are or may become available;
•our estimates of the number of patients that we will enroll in our clinical trials and the timing of their enrollment;
•the beneficial characteristics, safety, efficacy, and therapeutic effects of our drug candidates;
•our plans for further development of our drug candidates, including additional indications we may pursue;
•our ability to adequately protect and enforce our intellectual property and proprietary technology, including the scope of protection we are able to establish and maintain for intellectual property rights covering our current drug candidates and other drug candidates we may develop, receipt of patent protection, the extensions of existing patent terms where available, the validity of intellectual property rights held by third parties, the protection of our trade secrets, and our ability not to infringe, misappropriate or otherwise violate any third-party intellectual property rights;
•the impact of any intellectual property disputes and our ability to defend against claims of infringement, misappropriation, or other violations of intellectual property rights;
6
Table of Contents
•our ability to keep pace with new technological developments;
•our ability to utilize third-party open source software and cloud-based infrastructure, on which we are dependent;
•the adequacy of our insurance policies and the scope of their coverage;
•the potential impact of a pandemic, epidemic, or outbreak of an infectious disease, such as COVID-19, or natural disaster, global political instability, or warfare, and the effect of such outbreak or natural disaster, global political instability, or warfare on our business and financial results;
•our ability to maintain our technical operations infrastructure to avoid errors, delays, or cybersecurity breaches;
•our continued reliance on third parties to conduct additional clinical trials of our drug candidates, and for the manufacture of our drug candidates for preclinical studies and clinical trials;
•our ability to obtain, and negotiate favorable terms of, any collaboration, licensing or other arrangements that may be necessary or desirable to research, develop, manufacture, or commercialize our platform and drug candidates;
•the pricing and reimbursement of our current drug candidates and other drug candidates we may develop, if approved;
•our estimates regarding expenses, future revenue, capital requirements, and need for additional financing;
•our financial performance;
•the period over which we estimate our existing cash and cash equivalents will be sufficient to fund our future operating expenses and capital expenditure requirements;
•our ability to raise substantial additional funding;
•the impact of current and future laws and regulations, and our ability to comply with all regulations that we are, or may become, subject to;
•the need to hire additional personnel and our ability to attract and retain such personnel;
•the impact of any current or future litigation, which may arise during the ordinary course of business and be costly to defend;
•our expectations regarding the period during which we will qualify as an emerging growth company under the JOBS Act;
•our anticipated use of our existing resources and the net proceeds from our initial public offering; and
•other risks and uncertainties, including those listed in the section titled “Risk Factors.”
We have based these forward-looking statements largely on our current expectations and projections about our business, the industry in which we operate, and financial trends that we believe may affect our business, financial condition, results of operations, and prospects. These forward-looking statements are not guarantees of future performance or development. These statements speak only as of the date of this report and are subject to a number of risks, uncertainties and assumptions described in the section titled “Risk Factors” and elsewhere in this report. Because forward-looking statements are inherently subject to risks and uncertainties, some of which cannot be predicted or quantified, you should not rely on these forward-looking statements as predictions of future events. The events and circumstances reflected in our forward-looking statements may not be achieved or occur and actual results could differ materially from those projected in the forward-looking statements. Except as required by applicable law, we undertake no obligation to update or revise any forward-looking statements contained herein, whether as a result of any new information, future events, or otherwise.
7
Table of Contents
In addition, statements that “we believe” and similar statements reflect our beliefs and opinions on the relevant subject. These statements are based upon information available to us as of the date of this report. While we believe such information forms a reasonable basis for such statements, the information may be limited or incomplete, and our statements should not be read to indicate that we have conducted an exhaustive inquiry into, or review of, all potentially available relevant information. These statements are inherently uncertain and you are cautioned not to unduly rely upon them.
8
Table of Contents
Item 1. Business.
Overview
We are a clinical-stage biotechnology company industrializing drug discovery by decoding biology. Central to our mission is the Recursion Operating System (OS), a platform built across diverse technologies that enables us to map and navigate hundreds of billions of biological and chemical relationships within one of the world’s largest proprietary biological and chemical datasets, the Recursion Data Universe. Scaled ‘wet-lab’ biology and chemistry tools are organized into an iterative loop with ‘dry-lab’ computational tools to rapidly translate map-based hypotheses into validated insights and novel chemistry, unconstrained by published literature or human bias. Our focus on novel technologies spanning target discovery through translation, as well as our ability to rapidly iterate between wet lab and dry lab in-house and at scale, differentiates us from other companies in our space. Further, our balanced team of life scientists and computational and technical experts creates an environment where empirical data, statistical rigor and creative thinking are brought to bear on our decisions. To date, we have leveraged our Recursion OS to enable three value drivers: i) an expansive pipeline of internally-developed programs, including several clinical-stage assets, focused on genetically-driven rare diseases and oncology with significant unmet need and market opportunities in some cases expected to be in excess of $1 billion in annual sales; ii) strategic partnerships with leading biopharma companies to map and navigate intractable areas of biology, including fibrosis with Bayer and neuroscience with Roche and Genentech, to identify novel targets and translate potential new medicines to resource-heavy clinical development overseen by our partners; and iii) Induction Labs, a growth engine created to explore new extensions of the Recursion OS both within and beyond therapeutics. We are a biotechnology company scaling more like a technology company.
Figure 1.The Recursion Operating System (OS) for industrializing drug discovery. The Recursion OS is an integrated, multi-faceted system for iteratively mapping and navigating large-scale and rich biological and chemical datasets to industrialize drug discovery and translation.
The Digital Biology Opportunity
The traditional drug discovery and development process is characterized by substantial financial risks, with increasing and long-term capital outlays for development programs that often fail to reach patients as marketed products. Historically, it has taken over ten years and an average capitalized R&D cost of approximately $2 billion per approved medicine to move a drug discovery project from early discovery to an approved therapeutic. Such productivity outcomes have culminated in an industry success rate of 8% to 14% from discovery to commercialization, respectively, yielding a rapidly declining IRR for the industry, from 10% in 2010 to 2.5% in 2020.1-5
10
Table of Contents
Figure 2. Historical biopharma industry R&D metrics. The primary driver of the cost to discover and develop a new medicine is clinical failure. Less than 4% of drug discovery programs that are initiated result in an approved therapeutic, resulting in a risk-adjusted cost of $1.8 to $2.6 billion per new drug launched.1-51,2,3,4,5
These sobering metrics, despite incredible investment and brilliant scientists, point to the need to evolve a more efficient drug discovery process and explore new tools. Traditional drug discovery relies on basic research discoveries from the scientific community for disease-relevant pathways and targets to interrogate. Coupled with biology’s incredible complexity, this approach has forced the industry to rely on reductionist hypotheses of the critical drivers of complex diseases, which can create a ‘herd mentality’ as multiple parties chase a limited number of therapeutic targets. The situation has been exacerbated by normal human bias (e.g., confirmation bias and sunk-cost fallacy). Accentuating this problem, the sequential nature of current drug discovery activities and the challenges with aggregation and relatability of data across projects, teams and departments lead to frequent replication of work and long timelines to discharge the scientific risk of such hypotheses. Despite decades of accumulated knowledge, the result is that drug discovery has unintentionally become almost artisanal, creating major hurdles for innovation.
Contemporaneously, technological innovations, such as machine learning (ML) have transformed complex industries - from media to transportation to e-commerce - through the creation of scalable and continuously improving iterative cycles of digitization, data aggregation and prediction. The biopharma sector, however, has been slower to embrace such innovations and methods of thinking, except in very narrow areas. We are focused on filling this innovation gap by building a new type of drug discovery engine, the Recursion OS, and reengineering the end-to-end process from the ground up using multiple technological advances that have become accessible within just the past decade.
1Alacrita Consulting. Pharmaceutical Probability of Success. (2018)
2Deloitte. Ten years on: Measuring the return from pharmaceutical innovation (2020)
3DiMasi et al. Innovation in the pharmaceutical industry: New estimates of R&D costs. Journal of Health Economics. 47:20-33 (2016)
4Paul, et al. How to improve R&D productivity: the pharmaceutical industry’s grand challenge. Nature Reviews Drug Discovery. 9: 203-214 (2010)
5 Martin et al. Clinical trial cycle times continue to increase despite industry efforts. Nature Reviews Drug Discovery. 16:157 (2017)
11
Table of Contents
Figure 3. The standard iterative loop involves 1) profiling of real systems, 2) aggregation and analysis and 3) algorithmic inference, as used by machine-learning native companies across multiple industries6. The details of the real system that is profiled change based on the industry. For example, using satellite and street data along with traffic, construction and weather data to model the real world and predict optimal routes and points of interest along a route or the use of detailed user metrics from media viewing apps to map human preferences and predict and refine new content. Or in the entertainment context, using detailed measurement of viewing preferences to predict the most appealing media types. In these cases, and many others, digital maps of reality create ever-improving predictions that can be tested, leading to both a data moat and ever-improving products.
Our Radical New Approach to Drug Discovery
The emergence of technological innovations has created the opportunity to envision new approaches to discovering therapeutics at scale. We are pioneering the integration of these technological innovations across biology, chemistry, automation, data science and engineering to modernize drug discovery. Combining advances in high content microscopy with arrayed CRISPR genome editing techniques, we can rigorously generate massive, high-dimensional biological and chemical datasets to probe genome-scale biological contexts in multiple human cellular conditions, giving rise to the Recursion Data Universe. Simultaneously, exponential improvements in compute speed and reductions in data storage costs driven by the technology industry, married with ML tools to make sense of complex data, enable us to efficiently harness these massive datasets and perform an unbiased inquiry of causative human biology, unconstrained by presumptive hypotheses. We believe this will enable us to derive novel biological insights previously inaccessible to scientific researchers, reduce the effects of human bias inherent in discovery biology and reduce translational risk at the program outset. For example, given any gene of interest, our platform reveals its relationship to all other genes and molecules included in the Recursion Data Universe, based on proprietary data created in our own automated wet laboratory. Thus we are vastly expanding the scope of surveyable biology and combining novel, basic science and therapeutic discovery into a single step.
6Adapted from “Around the physical-digital-physical loop - A current look at Industry 4.0 capabilities” Deloitte Insights 10 October 2018, Rutgers and Sniderman.
12
Table of Contents
Figure 4. The productionized portions of the Recursion OS today. We use our proprietary software and highly-automated wet laboratory to design and execute up to 2.2 million experiments each week across diverse biological and chemical matter. Complex, high-dimensional data from these experiments are generated at a rate of up to 110 terabytes per week and aggregated and analyzed by proprietary neural networks in either distributed cloud computing environments or on our own high-performance compute cluster, BioHive-1. We leverage these algorithms to make predictions about the relationships between untested combinations of biology and chemistry. As of today, we have made more than 200 billion such predictions. Our scientists navigate this vast Map of Biology using proprietary software to discover novel relationships, which we can quickly test either in-house across a variety of assays or via clinical research organizations (CROs). As we validate or refute the predictions in orthogonal assays, up to and including complex animal models, our Recursion OS is continuously improved. This iterative cycle of mapping and navigating is akin to the strategy used by many of the largest technology companies in other complex industries.
13
Table of Contents
Figure 5. Our radically new approach to drug discovery. To date, we have used our approach to generate one of the largest biological and chemical datasets on earth, at nearly 13 petabytes, which is growing by up to 2.2 million experiments’ worth of data each week. In addition, we have built a proprietary suite of software applications within the Recursion OS, making us well-positioned to automate and accelerate basic science and drug discovery tasks and enable scientific teams to quickly and iteratively evaluate therapeutic candidates. Cumulatively, these advances may redefine R&D productivity, as technology has disrupted many other industries, and we believe they will generate forward program growth as they have led to forward revenue growth in the context of technology companies. By applying the Recursion OS to drug discovery, Recursion expects to turn drug discovery from sequential trial-and-error into a search problem where we map and navigate biology in an unbiased manner to discover new insights and translate them into potential new medicines at scale.
Recursion: A Biotechnology Company Scaling More Like a Technology Company
Traditional approaches to drug discovery typically begin with a specific indication and a human-derived target hypothesis. Bespoke assays are subsequently built, and data is generated to identify therapeutic candidates acting against the proposed target. In contrast, we empirically generate large datasets encompassing a broad range of indications, with data across hundreds of thousands of biological and chemical perturbations. We combine this data within our Recursion Data Universe, with the proprietary suite of advanced computational tools in our Recursion OS to map the relationships among and between all of the possible combinations of perturbations. We then initiate and advance new therapeutic programs by navigating the map to the most exciting predicted relationships. Mutually reinforcing advances in ML algorithms and an ever-growing body of knowledge through continuous data generation create a flywheel of novel insights, increasing the efficiency and output of our pipeline. Further, the time and cost for us to explore a hypothesis are radically less than traditional methods and approaches require, meaning we can explore biology and chemistry much more broadly to find the best relationships for translational research.
14
Table of Contents
Table 1. The scale and acceleration of our growth along multiple axes. We are a biotechnology company scaling more like a technology company, as demonstrated by our growth in inputs (experiments) and growth in outputs (data, biological and chemical relationships, programs and partnerships). (1) Includes approximately 500,000 compounds from Bayer’s proprietary library. (2) ‘Predicted Relationships’ refers to the number of Unique Perturbations that have been predicted using our maps. (3) Announced a collaboration with Roche and Genentech in December 2021 and received an upfront payment of $150 million in January 2022.
The Recursion OS
Using our highly-automated wet-lab infrastructure, we have executed approximately 115 million experiments across different biological and chemical contexts in multiple human cell types. The resultant Recursion Data Universe, which grows nearly constantly as new experiments are performed, is the substrate by which we use sophisticated computational techniques to Map the underlying biology and chemistry. We apply additional sophisticated computational techniques to these Maps to build our Navigating Tools, which allow us to predict hundreds of billions of biological and chemical relationships in silico and prioritize the most novel and promising candidates for further validation in our wet laboratories. Our mapping and navigating approach to drug discovery means that the ambitious experimental explorations that would have taken us over 1,000 years to execute physically can now be inferred in a matter of months due to the relatability of the dataset that we have already constructed. To date, we have built, validated and deployed our approach with a focus on novel target discovery and validation, which we view as the most challenging step in the drug discovery process due to the bias and limitations of the modern reductionist approach to discovery. We continue to invest in extending our approach into chemistry to enable us to act more rapidly and with higher success rates in translating our novel target discovery work into IND-enabled programs. In the future, we expect that we will further evolve our approach into techniques that improve our ability to execute clinical programs at scale. Though still early, we believe we have demonstrated meaningful leading indicators that our approach industrializes drug discovery, broadening the funnel of potential therapeutic starting points, identifying failures earlier in the research cycle when they are relatively inexpensive and accelerating the delivery of high potential drug candidates to the clinic while reducing cost.
15
Table of Contents
Figure 6. The Recursion OS today, along with a roadmap for future extensions and evolution. In its ideal state, a drug discovery funnel would be shaped like the letter ‘T,’ where a broad universe of possible therapeutics could be narrowed immediately to the best candidate, which would advance through subsequent steps of the process quickly and with no attrition. Our goal is to leverage technology to reshape the typical drug discovery funnel towards its ideal state by rapidly narrowing the funnel. Late-stage clinical failures are the primary driver of costs in today’s pharmaceutical R&D model, due in part to inherent uncertainty in the clinical development and regulatory process. Reducing the rate of costly, late-stage failures and accelerating the timeline from hit to a clinical candidate would create a more sustainable R&D model.
Figure 7. Reshaping the drug discovery funnel. The aim of the Recursion OS is reshaping the traditional pharma pipeline into a more ideal funnel in which the broad swath of biological and chemical data fed into the platform are quickly triaged and fed into an accelerated translation path into the clinic.
We believe we have made progress in reshaping the traditional drug discovery funnel in the following ways:
16
Table of Contents
•Broaden the funnel of therapeutic starting points. Our flexible and scalable Mapping Tools and Infrastructure enable us to infer hundreds of billions of relationships between disease models and therapeutic candidates, ‘widening the neck’ of the discovery funnel beyond hypothesized and therefore human-biased targets.
•Identify failures earlier when they are relatively inexpensive. Our proprietary Navigation Tools enable us to explore our massive biological and chemical datasets to validate more and varied hypotheses rapidly. While this strategy results in an increase in early stage attrition, we are able to rapidly prioritize programs with a higher likelihood of downstream success. Over time, and as our OS improves, we expect that moving failure earlier in the pipeline will result in an overall lower cost of drug development.
•Accelerate delivery of high-potential drug candidates to the clinic. Additionally, the Recursion OS contains a suite of digital chemistry tools that enable highly efficient exploration of chemical space, including 3D virtual screening as well as translational tools that improve the robustness and utility of in vivo studies.
We have leveraged our evolving Recursion OS to explore more than 150 disease programs to a depth sufficient to quantify improvements in the time, cost and anticipated likelihoods of program success by discovery stage compared to the traditional drug discovery paradigm. These metrics are leading indicators that, using our approach, we may be able to industrialize drug discovery. We believe that future iterations of the Recursion OS will enable even greater improvements. Ultimately, we look to minimize the total dollar-weighted failure while maximizing the likelihood of success.
Figure 8. The trajectory of our drug discovery funnel mirrors the ‘ideal’ pharmaceutical drug discovery funnel. We believe that, compared to industry averages, our approach allows us to: i) identify low-viability programs earlier in the research cycle, which quickly narrows the funnel, ii) spend less per program and iii) rapidly advance programs to a validated lead. Data shown are the averages of all our programs from 2017 through 2021. As we continue to evolve and expand our Recursion OS through improvements in chemistry, digital chemistry and predictive ADMET, we believe we will further improve overall R&D productivity.
Over time, we believe continued successes and improvements in any or all of the dimensions highlighted above will improve overall R&D productivity, allowing us to address targeted patient populations that may otherwise not be commercially viable using traditional drug discovery approaches. Further, we believe our unbiased approach may lead to novel targets and allow us to outperform others in highly competitive disease areas where multiple parties often simultaneously pursue a limited number of similar target hypotheses. These advantages potentially significantly expand the total addressable market for our technology. However, the process of clinical development is inherently uncertain, and there can be no guarantee that we will achieve shorter development timelines with future product candidates.
Our Business Strategy
17
Table of Contents
Figure 9. We harness the value and scale of our maps of biology using a capital efficient business strategy. Our business strategy is segmented into our: i) internal pipeline focused on oncology, rare diseases and other capital efficient opportunities, ii) enterprise-scale discovery partnership agreements in large therapeutic areas such as fibrosis with Bayer and neuroscience with Roche and Genentech and iii) Induction Labs which is our growth engine for translating our platform into auxiliary business opportunities over a longer-term horizon (not depicted above).
Our business strategy is to build, explore and develop opportunities that we feel we are most uniquely suited to advance. While most biopharma companies are focused on a narrow slice of biology or therapeutic area, where they believe they have an advantage or insight, our vision is to decode biology by mapping and navigating broad and diverse datasets so that we can, over time, evolve and extend our Recursion OS to deliver valuable and translatable insights at scale and across many therapeutic areas and modalities. Success in this endeavor would create extraordinary value and impact. Today, the biopharma industry has a market capitalization of multiple trillions of dollars, and creates products that touch nearly every human in the world at some point in time. Yet, on average, products developed in our industry fail in clinical development 90% of the time. This industry-wide inefficiency means continued investment in refining our Recursion OS to improve the probability of success of our programs over time is by far the most valuable long-term driver of our success, and it is also what we are most uniquely positioned to deliver.
Delivery of subsequent iterations of the Recursion OS, however, requires that we make tangible demonstrations of progress and potential along the way. As such, we developed a multi-pronged, capital efficient business model focused on three key value-drivers that enable us to demonstrate our progress over time while continuing to invest in the development of the Recursion OS, which we are convinced is our most compelling long-term value driver.
Value-Driver 1 - Near-Term Wholly Owned, Capital Efficient Programs
We believe that the primary currency of any biotechnology company today is clinical-stage, wholly-owned assets. These programs can be concretely valued using a variety of models by key stakeholders in the biopharma ecosystem and present the potential to meet critical patient needs. Further, for Recursion, these assets have a variety of additional benefits, including: a) validation of key elements of the Recursion OS; b) growing our expertise in clinical development; and c) building in-house processes to interact with regulatory agencies and advance medicines towards the market. This last point is perhaps the most important for Recursion. If the Recursion OS evolves in the manner we have designed it to, it will improve with more iterations such that future programs could be more valuable than today’s programs. In this way, operating as a vertically-integrated biopharma company that leverages technology at every step from target discovery through clinical development (and even marketing and distribution) may be the long-term business model with the most upside for our stakeholders, including both investors and patients. Thus, the importance of our early cycles of learning and iteration in clinical development
18
Table of Contents
have long-term value that may exceed the near-term commercial opportunities of any of the indications we have chosen to explore. For these reasons, we have directed our internal programs in areas that are both diverse and capital-efficient. Moreover, we may be opportunistic about selling or licensing assets after they achieve key value-inflection milestones so that we can re-invest in our long-term strategy.
Value-Driver 2 - Intermediate Term Partnered Programs
We believe that in its current form, our Recursion OS is already capable of delivering many more therapeutic insights than we would be able to responsibly shepherd alone today. As such, we have chosen to partner with experienced, top-tier biopharma companies to explore intractable and resource-intensive areas of biology like fibrosis with Bayer and neuroscience with Roche and Genentech. The key advantages of these partnerships are that: i) we are able to deploy the Recursion OS to turn latent value into tangible value in areas of biology where it would be challenging for us to do so alone; ii) the clinical development paths for these large therapeutic areas are often resource-intensive and highly complex; and iii) we are able to learn from our colleagues at these top-tier companies such that it could give us a competitive advantage in the industry over the longer term. This strategy also embeds us in the discovery process of large pharmaceutical companies, and gives rise to an alternative long-term business model whereby we become a valued partner of many such companies, focusing on discovery and de-risking of broad and varied programs while relying on our partners to develop and market the medicines while we take an increasingly large portion of the upside. Based on how value is ascribed across our industry today, this model alone is not yet feasible to maximize our business impact. However, we feel that shifts in industry perception and improving economics associated with each partnership agreement that we sign suggest that there is some potential for this portion of our business model to become the most value-accretive over the long-term.
Value-Driver 3 - Induction Labs for Long-Term Value Impact
Mapping and navigating biology has extraordinary potential to create better medicines faster and at lower costs. This is our primary focus today and is likely to be the most impactful use of our Recursion OS. However, there may be tangential markets and opportunities in spaces like diagnostics for which the infrastructure and technology we have built could create compelling value, impact and operating synergies. We will continue to make very small exploratory investments to test the utility of our platform to create new value-drivers for Recursion over the longer term.
Our Programs
Every program at Recursion is a product of our Recursion OS. All of the programs in our internal pipeline are built on unique biological insights surfaced through the Recursion OS and target diseases where: i) the disease-causing biology is well defined, but the downstream effects of the disease-cause are typically poorly understood or where the primary targets are typically considered undruggable and ii) there is a high unmet medical need, there are no approved therapies or there are significant limitations to existing treatments. Several of our internal pipeline programs target indications with market opportunities expected to be near to or in excess of $1.0 billion in annual sales and we are preparing for three programs to enter Phase 2 or Phase 2/3 clinical trials within the first three quarters of 2022 and a fourth program to enter a Phase 1 clinical trial within the second half of 2022.
19
Table of Contents
Figure 10. Examples of current Recursion programs falling into our First, Second and Next Generation paradigms. The earliest iterations of the Recursion OS leveraged brute-force search (where small molecules were tested directly in the context of each disease model we built) and used a small molecule library restricted primarily to known chemical entities. Programs arising from this iteration of the Recursion OS are deemed First Generation Programs. As we developed our chemistry capabilities and new chemical entity library at Recursion, Second Generation Programs arose, though the throughput needed to screen large libraries of new chemical entities presents a powerful but relatively inefficient solution. Today, most of our new programs, as well as new partnerships or expansions of prior partnerships, are Next Generation Programs, whereby we use our maps of biology to navigate to novel or unexpected relationships between molecules (known or new chemical entities) and then validate those predictions in our wet labs.
•Recursion’s First Generation of Potential Medicines. The following programs represent the novel use of a known chemical entity discovered using early iterations of the Recursion OS.
◦REC-994 for the treatment of cerebral cavernous malformation, or CCM— Phase 2a enrolling patients at the time of filing. Orphan Drug Designation granted in the US and EU.
◦REC-2282 for the treatment of neurofibromatosis type 2, or NF2—expected Phase 2/3 initiation in Q2 2022. Orphan Drug Designation in the US and EU, as well as Fast-Track Designation in the US, have been granted.
◦REC-4881 for the treatment of familial adenomatous polyposis, or FAP—expected Phase 2 initiation in Q3 2022. Orphan Drug Designation granted in the US.
◦REC-3599 for the treatment of GM2 gangliosidosis, or GM2—expected Phase 2 initiation in 2024.
20
Table of Contents
•Recursion’s Second Generation of Potential Medicines. The following programs arose from a brute-force approach leveraging either an expanded internal new chemical entity library or a partner new chemical entity library.
◦REC-3964 for the treatment of C. difficile colitis— expected Phase 1 initiation in 2H, 2022
◦REC-64917 for Neural or Systemic Inflammation
◦Multiple simultaneous programs in fibrosis advancing with Bayer
•Recursion’s Next Generation of Potential Medicines. The following programs represent a promising subset of known or new chemical entities discovered and developed using the latest Recursion OS mapping and navigating tools.
◦REC-65029 and derivatives or functionally related series for the Treatment of HRD-negative Ovarian Cancer by leveraging a potentially novel target insight
◦REC-648918 and derivatives or functionally related series to enhance anti-tumor immune response leveraging a potentially novel target insight (Target Alpha)
◦REC-2029 for the treatment of Wnt-mutant Hepatocellular carcinoma
◦REC-14221 and derivatives or functionally related series for the treatment of solid and hematological malignancies using indirect MYC inhibition
◦REC-64151 and derivatives or functionally related series for the treatment of immune checkpoint resistance in KRAS/STK11 mutant non-small cell lung cancer
◦Potential future programs in fibrosis with Bayer or in neuroscience or a single oncology indication with Roche and Genentech
In addition to the programs highlighted above, we are actively developing dozens of additional programs which may prove to be drivers of our future growth. As we have significantly expanded our chemistry capabilities in the last year and continue to invest deeply in these key elements of the Recursion OS, moving forward we expect that the vast majority of our new programs will be part of our Next Generation of potential programs discovered using our tools for mapping and navigating biology. We believe that the number of potential programs we can generate with our Recursion OS is key to the future of our company, as a greater volume of validated programs has a higher likelihood of creating value. The speed at which our OS generates a large number of product candidates is important, since traditional drug development often takes a decade or more. In addition, we believe that our large number of potential programs makes us an attractive partner for larger pharmaceutical companies. The static or declining level of R&D output at many large companies means that they have an ongoing need for new projects to fill their pipelines.
21
Table of Contents
Figure 11. The power of our Recursion OS as exemplified by the breadth of active research and development programs. We have an expansive pipeline of internally-developed programs spanning multiple therapeutic areas and consisting of both new uses for existing compounds and new chemical entities, or NCEs, under active research and development. All populations are US and EU5 incidence unless otherwise noted. EU5 is defined as France, Germany, Italy, Spain and the UK. (1) Prevalence for hereditary and sporadic symptomatic population. (2) Annual US and EU5 incidence for all NF2-driven meningiomas. (3) Worldwide prevalence; conducting dose optimization study in animal model with a potential trial start in 2024 (4) US and EU5 prevalence (5) Our program has the potential to address a number of indications with systemic or neural inflammatory components. We have not finalized a target product profile for a specific indication. (6) Our program has the potential to address a number of indications driven by MYC alterations, totaling 54,000 patients in the US and EU5 annually. We have not finalized a target product profile for a specific indication. (7) Our program has the potential to address a number of indications in this space.
Our People
While we operate at the intersection of cutting-edge science and technology from multiple disciplines, our people are the glue that holds us together and are the most important part of our company. Unlike traditional biotechnology companies, our rapidly growing team of approximately 400 Recursionauts is balanced between life scientists such as chemists and biologists (approximately 40% of employees) and computational and technical experts such as data scientists and software engineers (approximately 35% of employees), creating an environment where empirical data, statistical rigor and creative thinking is brought to bear on the problems we address. While we are united in a common mission, Decoding Biology to Radically Improve Lives, our greatest strength lies in our differences: expertise, gender, race, disciplines, experience and perspectives. Deliberately building and cultivating this culture is critical to achieving our audacious goals. Read more about how we invest in and motivate our people to achieve our mission in Recursion’s first Environmental, Social and Governance Report, released simultaneously with this annual report.
The Recursion OS - In Depth
The Recursion OS is an integrated, multi-faceted system for iteratively mapping and navigating massive biological and chemical datasets to industrialize drug discovery. It consists of three parts:
•Mapping Tools and Infrastructure: A synchronized network of highly scalable enabling hardware and software used to design and execute diverse biological experiments and subsequently store our ever-growing datasets. One of the cornerstones of this layer is our state-of-the-art ML supercomputer, BioHive-1, which we believe is one of the most powerful supercomputers wholly owned by any single biopharma
22
Table of Contents
company for drug discovery applications and within the top 100 most powerful supercomputers across any industry.
•The Recursion Data Universe: As of December 31, 2021, our Recursion Data Universe contained nearly 13 petabytes of highly relatable biological and chemical data spanning phenomics, orthogonomics, InVivomics and bespoke bioassay data.
•Navigating Tools: A suite of in-house software tools, algorithms and machine learning approaches designed to explore data from the Recursion Data Universe and translate it into actionable insights for our research and development teams.
The combination of wet-lab biology and dry-lab computational tools are organized in an iterative loop to rapidly translate map-based hypotheses into validated insights and novel chemistry, unconstrained by published literature or human bias. While many in the industry have focused on point-solutions and digital chemistry tools, our focus on novel technologies spanning target discovery through translation as well as our ability to rapidly iterate between wet lab and dry lab differentiates us from other companies. More importantly, our repetition of wet-lab validation and in silico predictions creates a flywheel effect, where data generation and learning accelerate side-by-side and further strengthen our drug discovery platform. While emerging competitors and large, well-resourced incumbents may pursue a similar strategy, we have two advantages as a first mover: i) no amount of resources can compress the time it takes to observe naturally occurring biological processes, and ii) the ever-growing Recursion Data Universe creates compounding network effects that may make it difficult for others to close the competitive gap.
Figure 12: The Recursion OS for industrializing drug discovery. The Recursion OS is an integrated, multi-faceted system for iteratively mapping and navigating massive biological and chemical datasets to industrialize drug discovery. It is composed of: (i) Mapping Tools and Infrastructure, (ii) the Recursion Data Universe, which houses
23
Table of Contents
our diverse and expansive datasets and (iii) Navigating Tools, a suite of our proprietary discovery, design and development tools.
Mapping Tools and Infrastructure
Figure 13. Our Mapping Tools and Infrastructure generate our proprietary data. This layer is the backbone upon which the Recursion OS operates and comprises diverse and highly advanced enabling hardware and software systems working in concert.
The foundational layer of the Recursion OS is a highly-synchronized network of enabling hardware and software used to design, execute, aggregate and store the nearly 13 petabytes of rapidly growing, biological and chemical data. Discrete components of this layer include the following:
Biological Tools
We deliberately designed our platform to model a wide range of biology spanning multiple therapeutic areas, including oncology, immunology, neuroscience, cardiovascular, metabolic and infectious diseases using the same, image-based endpoint and core technology stack. Our modular design enables us to systematically expand our search space into new areas of exploration while minimizing the need for bespoke assay development. In subsequent steps of our process, our modular design and consistent protocol enable us to analyze and compare the resulting data across these modules, revealing the interconnectedness of human biology and tractable therapeutic starting points. Modules that comprise our biological tool suite include:
•Genetics Module: A set of proprietary protocols and whole-genome arrayed guide RNA library using CRISPR gene editing approaches to model gene deficiency of every gene in the human genome in an arrayed and high-throughput format, plus BacMam capabilities to model gain-of-function.
•Soluble Factor Module: Proprietary protocols using single soluble factors such as cytokines and chemokines, or combinations thereof, to model a broad range of immune-related and complex diseases.
24
Table of Contents
•Infectious Disease Module: Proprietary protocols using diverse biological pathogens driving a broad range of infectious diseases as well as agents involved in the innate immune response (e.g., LPS, cyclic dinucleotides, etc.).
•Fibrosis Module: Proprietary models and protocols developed in partnership with Bayer to study fibrotic diseases, including cell co-culture systems.
•Neural Module: Proprietary models and protocols developed in partnership with Roche and Genentech to study neuroscience diseases, including advanced genetic engineering methods and iPSC-derived human relevant models.
•Complex Multicellular Disease Tools: Advanced co-culture models to explore multifactorial diseases where cell-cell crosstalk is a critical driver of the disease states. These approaches are particularly relevant in immunology, where regulation between adaptive immune cells (i.e., T cells, B cells) and innate immune cells (i.e., monocytes, macrophages) is critical to understanding the full breadth of immunological responses.
•Patient-Derived Tools: Techniques to improve the translatability and speed at which we validate and translate early discoveries. We are actively sourcing patient cells (nearly 400 individual lines across more than 65 diseases sourced to date), reprogramming them to induced pluripotent stem cells, or iPSCs, and banking the resulting lines so that we can rapidly differentiate these cells into multiple tissue-specific states for downstream validation when needed.
We continue to build out additional biology tools and modules to further expand our search space, while maintaining a common, image-based endpoint to reduce complexity, increase flexibility and ensure the relatability of our ever-growing Data Universe. Over time, we plan to introduce additional variables such as variable imaging time points, 3D models and tissue-specific organoids that move our screens ever closer to human systems biology.
Chemistry Tools
Our in-house chemistry tools include physical compound collections, state-of-the-art compound storage and handling infrastructure, and high-precision analytical equipment. Our experienced team of chemists use this equipment, and a network of reputable CROs, to advance discovery efforts and deliver differentiated drug candidates.
We have access in-house to nearly one million small molecule starting points from a combination of commercial, semi-proprietary and proprietary sources and use this library to identify new chemical starting points for small molecule discovery campaigns. Approximately 500,000 of these compounds reside within the Recursion NCE library, curated by our medicinal chemists and designed for highly druggable chemical properties while avoiding undesirable chemical properties, such as poor solubility and permeability. While this library has been constructed to maximize chemical diversity, we have ensured that several analogs of many compound cores are included to help identify emergent structure activity relationships for early hits and enable rapid hit expansion into readily available analogs. Additionally, we have curated a selection of approximately 7,500 preclinical and clinical-stage compounds from public forums or filings, covering approximately 1,000 unique mechanisms, for which an abundance of existing data and annotations currently exist. Such molecules are frequently used as tools within our work and may be advanced as therapeutic programs if our maps reveal unique and previously undisclosed biological activity. Approximately another 500,000 compounds are from Bayer’s NCE library, for which we do not have structural information.
We plan to substantially increase the size and diversity of our NCE library over the coming years through a combination of partnerships and investments that are being made in our Closed Loop Automated Synthesis Suite (CLASS) which will eventually integrate sample management, synthesis and purification and in vitro ADMET and bioanalytical testing. We believe we have the potential in the next 3-5 years to meet or surpass the scale of large pharmaceutical company libraries that typically have between approximately 1.4 and 4 million compounds. Our next generation of wet-lab has been designed with the theoretical potential to store more than 60 million compounds onsite.
25
Table of Contents
Figure 14. Our internal chemical libraries are highly diverse. This visualization of the structural diversity of approximately 200,000 compounds from one of our small molecule NCE libraries, where compounds are clustered based on descriptors using t-distributed stochastic neighbor embedding, demonstrates the evenly distributed and diverse nature of our compounds. This diversity increases the probability that we capture useful biochemical interactions across a broad range of biology.
Mass Compound Storage & Handling. We have invested in a sophisticated compound management infrastructure that allows for the environmentally controlled (temperature and humidity) storage of over one million compounds in tubes and plates. Our system enables rapid creation of purpose-built and custom libraries from our existing compound inventory. In addition, automated pipetting systems are in place to consistently aliquot and dilute these compounds into a variety of configurations for experimentation. All key events and lab data are tracked in our laboratory information management software, which integrates with experiment design and scheduling software, enabling accurate and seamless information tracking for our experiments.
Medicinal Chemistry/CMC Outsourcing. Our internal team of experienced medicinal chemists execute all drug design activities in-house but outsource drug synthesis and select ADMET assays to a network of reputable contract research organizations (CROs) with whom we have built well-established relationships. This may change with the build out of our CLASS system in the coming years. However, today external CROs provide easily scalable and project-specific resource flexibility, access to diverse chemistry expertise and rapid turnaround as we iterate on SAR. As programs advance into more advanced preclinical stages where synthesis at scale is of higher priority, our medicinal chemists work ever-closer with our CROs and internal chemistry, manufacturing and controls (CMC) group to craft detailed material plans for preclinical, IND-enabling and clinical supplies. We are also investing in the design and build out of a manufacturing facility that will enable us to synthesize and scale drug substance manufacture to support preclinical animal studies and early human clinical trials.
Analytical and Bioanalytical Chemistry. We have built an analytical laboratory equipped with state-of-the-art liquid chromatography-mass spectrometry equipment. Our lab performs analytical work to assess compound purity and identification for quality controls, bioanalytical work measuring compound levels in plasma and tissue samples from in vivo ADME and efficacy studies and plasma protein binding and permeability studies. Furthermore, this team carries out biomarker identification and validation activities in support of preclinical and clinical translational efforts. Further, we are investing in the design and build out of our Closed Loop Automated Synthesis Suite (CLASS). As currently designed, this suite will eventually integrate, both physically and virtually, sample management, synthesis and purification and in vitro ADMET and bioanalytical testing. The suite will be built on our digitalization, analytics
26
Table of Contents
and informatics capabilities to create an integrated computational platform with visualization tools, in-silico predictive models and retrosynthetic intelligence when fully mature. This suite will allow us to industrialize the synthesis of small molecules and subsequent data generation at scale
Cell Culture
We have built a state-of-the-art cell culture facility to consistently produce high-quality, mammalian cells, such as vein, kidney, lung, liver, skin and blood cell subsets, that go into each experiment run on our platform and in subsequent validation. We utilize a toolbox of in vitro cell culture techniques to scale production while driving down costs. This includes the graduated use of small scale flasks, with a 25 cm2 growth surface area, to large-scale, single-use bioreactors, with a combined 575,000 cm2 growth surface area that enables us to generate 25 billion human cells, enough for up to 8,000, 1536-well plates for screening.
We have on-boarded innovations including large scale, microcarrier-based, suspension culture systems to reduce footprint and increase growth surface for additional scale. Additionally, our cell culture facility is now fully equipped to perform work using human induced pluripotent stem cell (iPSC) lines, including: (i) CRISPR genome editing technologies to generate knock-out or knock-in lines (ii) differentiation of human induced pluripotent stem cells at large scale and (iii) increased cryostorage capacity for pluripotent cell lines and differentiated cell products. We will continue to onboard additional cutting-edge innovations to scale our work further.
Table 2. Numerous and diverse cell types onboarded to our platform enable us to broadly interrogate biology. Approximately 40 human cell types have been onboarded to our high-throughput discovery systems to date, spanning primary cells, cell lines and cells derived from iPSCs.
We maintain a strong track record of quality and consistency in our cell culture facility by implementing facility design and control systems that are uncommon among technology-enabled drug discovery companies. These designs and controls include rigorous process validation and documentation, a personnel training and qualification program, and routine quality monitoring. Our quality system is designed such that we routinely monitor our performance to identify and implement the appropriate preventive and continuous improvement actions.
Lab Robotics
We have assembled and synchronized robotic components, such as liquid dispensers, plate washers and incubation stations, that enable us to efficiently execute up to 2.2 million experiments per week with only a small
27
Table of Contents
team overseeing the process at any given time. These robotic systems are modular by design and easily configurable to allow us to create complex and variable workflows. This flexibility is essential for executing experiments using our diverse biological tools (e.g., genetic and soluble factors) and chemical libraries at scale and with high quality.
We ensure our lab generates consistent, accurate and precise data through the use of multiple systems: facility controls to prevent contamination of cells, rigorous assay validation and instrument qualification to ensure consistency, and routine quality monitoring to capture data automatically and track all critical experiment specifications. Our quality system is designed such that we routinely monitor our performance to identify and implement the appropriate preventive and continuous improvement actions.
Figure 15. Our high-throughput automation platform looks more like a sophisticated manufacturing facility than a biology R&D laboratory. Our platform can execute up to 2.2 million experiments each week with high-quality to enable downstream analyses.
Figure 16. The automated workflow used to generate our large-scale, image-based dataset. A core dataset in the Recursion Data Universe is based on over one billion labeled images of human cells generated across millions of unique perturbations (i.e., gene knockout, soluble protein factor addition, drug addition or combinations thereof) generated in our own wet laboratories.
28
Table of Contents
Our laboratory operates approximately 50 weeks each year. Since 2017, we have at least doubled our throughput every year while meeting our quality benchmarks. We have achieved this level of operational excellence by integrating state-of-the-art technology and adopting lean manufacturing principles. Our move to mapping and navigating biology using inference, and away from brute-force screening, has relieved some of the demand for exponential scaling of our platform moving forward.
Figure 17. The experimental throughput of our high-dimensional phenomics assay has scaled significantly over time. The capabilities of our phenomics assay have grown throughout 2021 with quick recovery following a COVID-19-induced, full-office closure in early 2020.
Data Capture
The Recursion Data Universe contains nearly 13 petabytes of highly relatable biological and chemical data spanning multiple different “omic” modalities. We have invested in state-of-the-art equipment to capture this data at scale and processes to ensure that the highest quality data are fed into the Recursion Data Universe.
High-Throughput Microscopy. Central to the Recursion Data Universe is our image-based dataset. As of December 2021, we operate 20 ImageExpress microscopes in our labs, which we believe to be the greatest single number of such systems in a single facility anywhere. These microscopes run nearly continuously, capturing over 155,000 fluorescent microscopy images every hour across six imaging channels. Alerts are automatically triggered if quality issues are detected, enabling our teams to quickly reimage our experimental conditions to obtain higher quality data. Upon imaging, our digital data pipeline immediately uploads these images to the cloud where they are processed within seconds. On a weekly basis, our pipeline captures, uploads and processes up to 110 terabytes of imaging data to add to the Recursion Data Universe.
High-Throughput Sequencing. Our high-throughput sequencing system enables us to profile transcriptomic measurements in house for any cell type and biological perturbations we develop. As of December 2021, this system includes two Illumina NovaSeq 6000 production scale sequencers. These sequencers currently process 6,100 individual transcriptome samples per week in development operations, and we anticipate a full production capacity of 44,000 transcriptomes a week in the future. Additionally, we have installed a 10x Chromium X instrument capable of performing single cell RNAseq workflows and are currently in the process of validating this assay. The addition of the single cell RNAseq platform will allow us an additional level of granularity in assessing transcriptional changes not capable with other transcriptomic methods.
In Vivo Data Collection. We use our proprietary cage hardware and continuous, high-resolution video systems to collect InVivomics data at scale. In 2021, we had 19 cage systems operational, actively surveying a total of 931 possible simultaneous in vivo subjects undergoing pharmacokinetics, efficacy and safety studies of our drug candidates, as well as R&D studies to unlock future predictive assays. This data is uploaded to the cloud where it is
29
Table of Contents
automatically analyzed. Readouts are provided back to our scientists and integrated into our Recursion Data Universe.
Figure 18. Our proprietary, scalable Smart Housing System for in vivo studies automatically collects and analyzes video and sensor data from all cages continuously.
Additional Data Collection Systems. Beyond phenomics, orthogonomics and InVivomics, we continuously capture experimental data from bespoke assays as we validate our discovery programs. Example data capture infrastructure includes multiplexed readouts for biological analytes, flow cytometry and electric cell-substrate impedance testing. As this data is generated, it is included in our data warehousing system that connects one-off experimental assays with the rest of the Recursion Data Universe.
Technology Stack
The Recursion OS is built on top of a core technology stack that is highly scalable and flexible. We have adopted a ‘hybrid-cloud’ strategy, leveraging the benefits of both public and private cloud infrastructure depending on the context and our needs:
•Public Cloud. The public cloud is our default choice for production workloads and applications. The scale, elasticity of compute and storage and economies of scale offered by public cloud computing providers enable us to cost-effectively execute our strategy.
•Private Cloud. The private cloud, or edge computing, is used to integrate our lab data flows, including the upload of data to the public cloud.
•BioHive-1 and High Performance Computing in a Private Cloud. In December 2020, we made a significant investment to expand our computing power, purchasing a world-class supercomputer named BioHive-1. BioHive-1 is built on NVIDIA’s DGX SuperPod architecture and ranked 97th on the most recent TOP500 list of the world’s most powerful supercomputers as of November 2021. This new computing power allows us to iterate on new neural network architectures faster and more efficiently, accelerating our deep learning models and empowering our growing workforce of ML experts. Deep learning projects that took a week to run on our previous cluster can run in under a day on the new cluster.
30
Table of Contents
Figure 19. We believe BioHive-1 is one of the most powerful supercomputers dedicated wholly to drug discovery for a single company. BioHive-1 consists of 40 NVIDIA DGX A100 640GB nodes which further expands our capability to rapidly improve ML models.
Enabling Software Tools
Alongside our infrastructure, we have built a suite of tools that empower our scientists to accurately design, execute and verify the quality of up to 2.2 million diverse experiments each week, spanning phenomic, orthogonomic and ADMET assays. Our tools, which take into account real-time onsite reagent supplies, enable consistent control strategies and design standards that make each week’s data relatable across time. Additionally, these tools automatically flag experiments or processes which miss quality requirements or stall at some point in the process and notify the appropriate Recursionaut, providing them the tooling needed for manual intervention. Elements of our Enabling Software Tool suite include:
•Experiment Design Tools: Proprietary Laboratory Information Management System (LIMS) to track reagent inventory and flexibly select compounds from our library, custom applications used to design large experimental layouts consisting of millions of perturbation conditions with appropriate randomization and control strategies, and proprietary algorithms for designing CRISPR gene editing guide RNAs for maximal knockout efficiency
•Experiment Execution and QC Tools: Suite of tools and dashboards to automatically execute and continuously monitor experimental protocols to ensure reliable experiment execution and custom web applications that enable our scientists to view and interact with microscopy images and associated meta-data from our phenomics platform for systematic QC at both the image- and plate-level.
31
Table of Contents
Figure 20. Experiment Delight allows our biologists to design massive experiments while complying with our complex proprietary rules for layout. Experiment Delight is our internal experiment design tool used to rapidly create large-scale experiment sets with high flexibility, while integrating our proprietary rules for experiment layout learned over approximately a decade of iterative improvement. The graphical interface facilitates experiment plate layout specification.
The Recursion Data Universe
32
Table of Contents
Figure 21. The Recursion Data Universe is at the core of the Recursion OS. The central asset of the Recursion OS is the Recursion Data Universe, encompassing multiple data types that compound together, the whole providing greater insight than the sum of the parts.
The Recursion Data Universe comprises nearly 13 petabytes of highly relatable biological and chemical data, including: phenomics, orthogonomics, ADMET assays, InVivomics and bespoke bioassay data. These different data modalities are highly complementary as we advance drug discovery and development programs. Phenomic data provides a broad, foundational layer of biological and chemical data, while other datasets provide greater translational insights. The size of the Recursion Data Universe has nearly doubled in the last year.
33
Table of Contents
Figure 22. Diverse datasets within the Recursion Data Universe are highly complementary. The Recursion Data Universe consists of complementary datasets spanning multiple data modalities. While phenomics data can be generated cost-effectively and at scale, other datasets such as transcriptomics, proteomics and InVivomics offer increasing insight as we translate programs from early discovery through development.
Phenomics
At the core of the Recursion Data Universe is our proprietary cellular image dataset generated by our automated phenomics platform. While investigating various biological and chemical contexts, the readout remains constant: a fluorescent microscopy image that captures composite changes in cellular morphology; a cellular phenotype. We use our proprietary staining protocol to capture these changes in cellular morphology across nearly all of our phenomic experiments. This protocol, consisting of six subcellular dyes imaged in six different channels, has been optimized to capture a wide array of biology across nearly any human cell type that can be cultured and perturbed in laboratory conditions. As a result, we can capture the effects of a wide range of biological and pharmacological phenomena of interest, including phenotypic changes induced by small molecules, genetic gain- and loss-of-function, toxins, secreted factors, cytokines, or any combination of the above.
34
Table of Contents
Figure 23. Our fluorescent staining protocol images multiple large cellular structures to capture a holistic assessment of cellular state. We use fluorescent dyes to stain a set of common cellular substructures that are subsequently captured using fluorescent microscopy imaging. Combined with tools from the Recursion OS, this complex and rich biological data modality can inform a host of scientific questions. The top image is a composite of the 6 channels. It is followed by each of the 6 individual channel faux-colored images of HUVEC cells: nuclei in blue, endoplasmic reticula in green, actin in red, nucleoli in cyan, mitochondria in magenta and Golgi apparatus in yellow. The overlap in channel content is due in part to the lack of complete spectral separation between fluorescent stains.
Cellular morphology is a holistic measure of cellular state that integrates changes from underlying layers of cell biology, including gene expression, protein production and modification and cell signaling, into a single, powerful readout. Images are also two-to-four orders of magnitude more data-dense per dollar than other -omics datasets that focus on these more proximal readouts, enabling us to generate far more data per dollar spent to inform our drug discovery efforts. Indeed, since 2017 we have approximately doubled the capacity of our phenomics platform each year and currently generate up to 13.2 million images or 110 terabytes of new data to the Recursion Data Universe per week across up to 2.2 million experiments. Lastly, our phenomics approach builds on the recent explosion of powerful computer vision and ML approaches driven by the technology industry over the last half decade. Modern ML tools can be trained to identify the most salient features of images without relying on any pre-selected, disease-specific subject matter expertise, even if these features are imperceptible to the human eye. Using these tools, we can capture the aggregate cellular response induced by a disease-causing perturbation or therapeutic, and quantify these changes in an unbiased manner, freeing us from human bias. In contrast, traditional drug discovery relies on presumptive target hypotheses and bespoke biological signaling assays that only capture narrow, pre-determined biology and thus limit the scope of biological exploration.
Figure 24. ML algorithms can detect cellular phenotypes that are indistinguishable to the human eye. Most morphological differences within our images are too subtle for the human eye to detect, but ML algorithms like those we deploy in our Recursion OS can readily distinguish between them. The heatmap of similarities shown here between learned embeddings of these images shows clear separation of highly similar cellular changes while even well-trained cell biologists or pathologists would be hard-pressed to describe consistent differences between these cell cultures.
35
Table of Contents
Orthogonomics
Phenomics provides cost-effective, information-rich and functional biological data well-suited for broad biological exploration. However, other data modalities such as transcriptomics and proteomics can be highly complementary. Both of these approaches generate supplemental data that can be useful for i) unraveling the mechanism of action by which a compound is active and/or ii) more precisely measuring (and confirming) a compound’s functional activity and efficacy. While the costs to measure bio-molecules using these approaches are orders of magnitude more expensive compared to phenomics, this data can be highly informative in order to advance programs. In particular, when used in a targeted manner (e.g., to follow up on predicted potential mechanisms of action) rather than broad primary profiling, orthogonomic approaches may deliver net value even at a higher per-measurement cost. Additionally, if we are able to generate this data cost-effectively and at scale, we may be able to significantly reduce the time needed to develop specific assays on a per bio-molecule basis. Collectively, we refer to these alternative modalities as orthogonomics, the generation and integration of orthogonal -omics-level datasets as a part of the Recursion Data Universe.
Scaled Transcriptomics. We have developed an in-house laboratory process capable of profiling over 20,000 genes from samples drawn from any of our biological modules. Throughout 2021 we leveraged our transcriptomic data generation engine to accelerate our biological understanding of many of our programs. We currently have the capability of processing up to 6,100 individual transcriptome samples per week, and have generated 91,400 whole transcriptome observations as of the end of 2021. The incorporation of in-house production-scale sequencers has reduced our transcriptomics data turnaround time by 70%. We intend to continue to develop, mature and scale this technology as a means to obtain valuable orthogonal data and a deeper understanding of the biology and pharmacology of our programs and lead molecules.
Proteomics. In 2021, we executed thousands of screens of proteomic samples, obtaining proteoprints for over 7,000 proteins for each in vitro and in vivo sample studied, and leveraged this data across over a dozen internal programs to inform our research operating plans and obtain a deeper understanding of the biology and pharmacology of our programs and lead molecules.
Other Scaled -omics. Exploration and development of scaled metabolomics and lipidomics are on our roadmap as additional medium-throughput mechanisms for orthogonal validation.
ADMET Assays
While our phenomics platform has historically been used to identify signals of compound efficacy, we explored the use of our image-based readout to predict ADMET properties of promising compounds early in the drug discovery process. Poor in vivo pharmacokinetics, including unwanted side effects, are a major driver of late-stage drug program failures.
To train predictive ADMET models, our team has built large-format ADMET datasets spanning various compound liabilities including CYP inhibition, which can indicate a risk of complication from drug-drug interactions and hERG liabilities, which can suggest a heightened risk for heart arrhythmias. This ADMET data has been combined with phenomic and compound structure data to create early predictive models, winnowing those drug candidates with a higher likelihood of potential liabilities before investing time and resources.
InVivomics
In vivo studies are an important tool for providing an assessment of the efficacy and safety of a compound within the context of a complete, complex biological system. Similar to other steps within the drug discovery and development process, conventional in vivo studies are fraught with human bias and limited in the endpoints that they measure. Using our In Vivo Data Collection Infrastructure, we can collect more holistic measurements of an individual animal’s behavior and physiological state using continuous video feeds and our proprietary animal cages, surveilling animals in their home environment. By automating the process of data collection, we can amass uninterrupted data on animal behavior and physiology across days, weeks, or even months allowing for a more accurate and holistic assessment of the animal’s health state across the entirety of the study. This data can subsequently be used to create more abstract representations of animal behavior, potentially allowing us to rapidly phenotype new animal models and identify in vivo disease signatures that may be more relevant for assessing compound efficacy and potential liabilities.
36
Table of Contents
Bespoke Assays
In addition to the large format datasets described above, our team is experienced at developing custom assays needed for program-specific validation at a smaller scale. These assays encompass diverse biomolecules, including nucleic acids, proteins and lipids, allowing for complete coverage across diverse therapeutic areas. Representative examples of these bespoke assays include high-content protein translocation readers and multiplexed readers to measure protein changes, qPCR or bead-based technologies to measure panels of transcript changes, mass spectrometry to measure more challenging biomolecules, electric cell-substrate impedance sensing and flow cytometry to measure distinct cellular subpopulations.
As this data is generated, it is included in our data warehousing system that connects bespoke assays with the rest of the Recursion Data Universe.
Navigating Tools
Our Navigating Tools area rapidly growing suite of in-house software applications designed to process and translate data from the Recursion Data Universe into actionable insights for our research and development teams to accelerate programs.
Figure 25. Navigating Tools. Our Navigating Tools are a suite of proprietary data generation, discovery and development tools that explore and transform data into actionable insights. The combination of our proprietary data generation and software tools provides the basis for data-driven decision making.
Data Processing Tools
37
Table of Contents
To understand, explore and relate new or existing data in the Recursion Data Universe, we must normalize, transform and analyze the data. Our tools in this layer manage the streaming of our data at scale to the appropriate public and private cloud, the transformation of our images into mathematical representations through our in-house proprietary convolutional neural networks, and the standard and custom analyses performed on our data as parameterized and requested by users. Anomalies are flagged to the team for fast resolution.
Figure 26. We convert raw images into a list of features that allows cross-image comparison. Microscopy images are run through a deep convolutional network with an architecture similar to the one above. The network is trained on our phenomics data so that, layer by layer, each image is transformed into a list of 128 features representing the cellular biology in the image. The resulting features power downstream analysis.
Biological and Chemical Activity Assessment
Our activity assessment tools enable us to evaluate the robustness of diverse disease model phenotypes and subsequently measure the activity of potential therapeutic agents within these disease models. These tools are target-agnostic by design, explore cellular biology holistically and enable the exploration of many disease models and potential therapeutics simultaneously with no significant alteration to the core platform.
38
Table of Contents
Figure 27. Our proprietary user interface enables our biologists to rapidly identify compounds with maximum positive effect on a disease phenotype while minimizing side effects. The results from our empirical hit identification screens allow drug discovery teams to rapidly explore results and focus on compounds that are believed to be the most promising.
Program Insights
We translate processed data into actionable insights which fall into two categories: i) insights into underlying biology and early therapeutic starting points and ii) insights into the specific chemical substrate of interest. We mine the Recursion Data Universe to predict therapeutic activity and behavior that may seed new NCE programs or new uses for KCE programs. We use an additional suite of tools to infer a compound’s mechanism of action and potential ADMET liabilities based on measures of similarity to other high-dimensional landmarks in our dataset and predictive models incorporating images and chemical structure.
PhenoMap. PhenoMap is a massive relational database of biological and chemical perturbation phenotypes that allow us, based on phenotypic similarity, to infer the relationship between any two perturbations (or groups of perturbations) in silico. To date, we are able to infer over 200 billion biological and chemical relationships, which are generated solely by ML tools without any human bias and may allow us to understand the mechanisms underpinning disease and how to manipulate them. For example, we can query the similarity (or dissimilarity) created by the CRISPR-engineered knockout of any two genes from our whole-genome arrayed CRISPR screen, revealing both known and novel drug targets never before described in scientific literature. We can also query the similarity between any small molecule in our library and all genetic knockouts, uncovering a compound’s mechanism of action and, most importantly, can infer the activity of such molecules against high-value drug targets. Our ability to probe the relationships between any perturbation in our library (spanning the genome and approximately one million small molecules) changes drug discovery from an iterative trial-and-error process into a computationally driven search problem.
39
Table of Contents
Figure 28. The PhenoMap allows our team to simultaneously view multiple relationships between genes and compounds. Our PhenoMap enables us to rapidly explore inferred biological and chemical relationships in order to: i) discover targets, ii) predict active hits, iii) optimize for similar or dissimilar relationships, and iv) predict mechanisms of action.
We are looking to augment the above insights by including data and predictions related to physicochemical and structural information about compounds, synthesizable compounds not yet tested on our platform, ADMET assays, and in vivo experiments.
Compound Intelligence. Our Compound Intelligence (CI) tools generate early insights into specific therapeutic candidates, helping us to advance candidates with favorable properties while culling candidates with higher likelihoods of failure. Using one application of CI, we can elucidate the mechanism of action of NCE compounds either by comparing a compound’s phenotype to: i) those phenotypes from our whole-genome arrayed CRISPR experiments (querying whether the phenotype induced by inhibition of a small molecule mimics any genetic knockout in our library) or ii) those phenotypes induced by well-annotated compounds in our repurposing library. Using a different application within CI, we can use our growing ADMET dataset and computational models to predict specific ADME and toxicology endpoints for therapeutic candidates. Compounds with low predicted ADMET properties are advanced. Compounds with high predicted ADMET properties may be discarded or flagged for subsequent investigation.
Program Acceleration
Once insights have surfaced, our researchers have a suite of digital chemistry and translational tools at their disposal to optimize compounds and accelerate discovery and development programs.
Compound Atlas. Compound Atlas is a collection of our proprietary and commercially-available digital chemistry tools that enables our scientists to expand from promising therapeutic starting points into more diverse chemical structures using large, enumerated chemical libraries from vendors such as Enamine and WuXi. Scaffold Shopper, a module within Compound Atlas, can compare candidate compounds identified by our platform to over 12 billion ready-to-synthesize and off-the-shelf molecules based on our 3D chemical functionality and shape-based similarities within a matter of minutes and at a low computational expense. Additionally, we have built software that enables our chemists to rapidly assemble dense mini-libraries around reproducible and validated hit molecules to accelerate structure-activity relationship (SAR) establishment without requiring custom synthesis.
40
Table of Contents
Figure 29. Scaffold Shopper enables our chemists to rapidly identify read-to-synthesize and off-the-shelf compounds for hit expansion. Comparisons are based on 3D chemical functionality and shape-based similarities generated within a matter of minutes and at a low computational expense.
Molecular Firehose. Molecular Firehose filters the expansive search results from Compound Atlas, so that our medicinal chemists can rapidly prioritize molecules of interest. Chemists can dynamically filter search results with a range of molecular properties and both 2D and 3D-based similarity scoring to better identify an appropriate compound set to order for synthesis from our chemical vendors.
Figure 30. Molecular Firehose filters multiple properties to rapidly identify viable compounds to synthesize.
InVivomics Research Suite
41
Table of Contents
The InVivomics Research Suite is our proprietary collection of software tools that enable scientists to monitor and analyze behavioral and physiological data from ongoing and completed in vivo studies. Study data for individual animals or aggregated across study groups can be explored in near real-time, better ensuring that the final study data will be reproducible and interpretable. Continuous monitoring allows researchers to similarly flag unexpected effects that may arise from animal handling, dosing, or compound liabilities and modify or terminate the study as needed. At the end of the study, graphical and tabular data are automatically generated to aid in the evaluation of study results and the design of follow-up in vivo studies.
More importantly, continuous video feeds and our proprietary animal cages enable us to amass uninterrupted data on animal behavior and physiology across days, weeks, or even months. ML tools within our InVivomics Research Suite can then be used to create more abstract representations of animal behavior, allowing us to rapidly phenotype new animal models and identify in vivo disease signatures that may be more relevant for assessing potential compound safety and efficacy attributes.
Figure 31. InVivomics Research Suite allows our team to track and analyze a broad range of data in ongoing animal studies. These tools enable our in vivo scientists to monitor individual subjects through near real-time video feed and data generation and review study level data.
Data Warehousing System
We employ a data warehousing system that encompasses the Recursion Data Universe, electronic lab notebooks generated by our research scientists, and technical analyses posted to our internal knowledge repository by our data and ML scientists. This data warehousing system is centralized and accessible for authorized Recursionauts and helps preserve institutional knowledge, further collective learning, and generate ideas for new discovery and development tools.
Bridging from Recursion OS Insights to Program Advancement
Reason to Believe. In order to identify novel program starting points, it is critical that the Recursion OS can accurately predict relationships across diverse domains of biology. To confirm the accuracy of our predictions, we have demonstrated that our approach recapitulates hundreds of well-known biological pathways. In the example below, we illustrate our map based predictions for approximately 150 gene knockouts from canonical biological pathways and known agonists or antagonists of these same pathways. By comparing the phenotypes induced by these perturbations to one another using our Recursion OS, we observed that each perturbation creates a unique phenotype and phenotypes form clusters that recapitulate well-understood biological pathways, including genes involved in Bcl-2 signaling, NF-KB signaling, RAS signaling, JAK/STAT signaling, and TGFß signaling.
42
Table of Contents
Figure 32. Inferred relationships between genes and small molecules faithfully recapitulate well known biology. Above, we show a visualization of approximately 0.00001 % of our map of biology (~22,500 of 203 billion predicted relationships) produced by our Recursion OS for well studied genes and small molecules. Increasingly dark shades of red reflect an increasing degree of phenotypic similarity. Increasingly dark shades of blue reflect an increasing degree of phenotypic oppositeness or anti-similarity (which often suggest inhibitory relationships between genes, though possibly distal). Highlighted sections reveal expected relationships along well-studied biological pathways.
These findings not only validate the accuracy of our inference relationships, but also suggest that we can use our approach of mapping and navigating biology and chemistry to identify new drug targets or early therapeutic starting points to seed new drug discovery programs. While there is no typical drug discovery program, most programs proceed as follows.
Step 1: Navigate the Map to Identify Novel Biological Targets and/or Early Therapeutic Starting Points. Using the Recursion OS, we can profile, map, and subsequently navigate relationships among diverse biological perturbations, including CRISPR gene knockouts, soluble factors, bacterial toxins and small molecules based on the similarity (shades of red in the figure above) or anti-similarity (shades of blue in the figure above) of each perturbation’s high-dimensional phenotype. Using these relationships, we can elucidate both potential novel drug targets or early potential therapeutic compounds to start new drug discovery programs. With more than 200 billion predicted relationships, there are more potential programs in our maps of biology than we can prosecute. For example, at a ‘hack-week’ in 2021, more than 100 potential new drug discovery programs were elucidated by about a dozen teams over 7-10 days using these maps. The maps mean that generating a program hypothesis requires no new wet-lab work; scientists simply use our software tools to navigate biology and chemistry.
Step 2: Empirically Confirm Map-Based Insights. Having selected a target and/or compound of interest based on its inferred activity, we then physically screen candidate compounds in the disease-relevant background to confirm our predictions. These experiments, deemed ‘lightning screens,’ are designed to confirm predicted relationships of interest from our map within 1-4 weeks. Data from the direct confirmation of a map-based insight is funneled back to project teams who can then make go/no-go decisions on initiating a program.
Step 3:Orthogonally Validate Insights. In addition to understanding and refining the chemistry (see steps 4 and 5 below), project teams build research operating plans based on confirmed map-based insights. These plans span orthogonal in vitro validation of the relationship (e.g., using various cellular -omics technologies, such as transcriptomics), bespoke assay development and evaluation, animal model validation and/or patient-derived cellular assay evaluation. We strive for independence in our validation both in the disease models used (e.g., in vivo
43
Table of Contents
models, in vitro systems) and in the endpoints measured (e.g., phenomics vs. transcriptomics vs. tumor growth in an in vivo model).
Step 4: Predicting the Mechanism of Action. Having confirmed our predictions empirically, for NCE programs, our medicinal chemists work to further understand the mechanism by which compounds are operating using our maps, often in parallel with our orthogonal validation work in step 3 above. Our compound library contains approximately six thousand compounds with well-annotated mechanisms of action. Using our mapping and navigating software tools and the phenotypes from these compounds, as well as from thousands of genes that we have knocked out using our CRISPR-gene editing tools, our chemists can compare the phenotypes of our validated compounds to these high-dimensional landmarks and assess their degree of similarity to identify potential mechanisms of action.
Figure 33. Compounds with the same mechanism cluster together phenotypically. A UMAP plot where each dot represents a different compound. Compounds that are phenotypically similar reside closer together and recapitulate mechanistic similarities.
Step 5: Optimize Validated Compounds into Viable Drug Candidates. While a compound may be active in our screens, most early therapeutic starting points have low potency and undesirable drug properties and must be optimized before advancing into in vivo and, ultimately, human studies. During the lead optimization process, our chemists rely upon our phenomics platform to repeatedly measure changes in compound potency and selectivity that result from changes in compound structure. Our chemists also take advantage of our burgeoning suite of proprietary digital chemistry tools to conduct chemical expansion exercises across more than 12 billion molecules in our in silico library which we can then order for validation on the platform.
Because this process may extend over several months, it is critical that our platform assay is highly stable over time. To ensure this stability, we test that our assay can reproduce specific measures of compound activity, such as a compound’s EC50 (the concentration of a drug that gives half-maximal response) or max-effect (the maximal response), in experiments run weeks, or even months, apart.
44
Table of Contents
In the example below, we ran four separate experiments of a HIF2a inhibitor known to be active against our VHL disease model over a period of three months. Dose-response curves across all four runs demonstrate a high degree of overlap, including highly similar EC50s and max-effect. Our calculated minimum significance ratio from this study, a common industry metric of in vitro assay reproducibility over time, is 1.076, which is highly robust by industry benchmarks7. These results demonstrate the stability of our assay and the ability to use our phenomic platform as a basis for SAR.
Figure 34. Compound activity is reproducible across experimental runs. Dose response curves from multiple runs of the same tool compound against our disease model for VHL loss-of-function show high consistency with a minimum significance ratio of 1.076.
Step 6: Select and Advance Drug Candidates into Clinical Trials. After optimizing therapeutic drug candidates, we select those compounds that have the best chemical properties to advance through development and ultimately clinical trials. We have built the internal capabilities to drive clinical candidates through IND-enabling studies, regulatory approval processes, and human clinical studies. Collectively, members of our team have been involved in over 700 clinical programs, including recently completing our first SAD and MAD studies in 2019 and 2020, respectively. Additionally, we work closely with a team of external consultants across regulatory, CMC, and clinical operations to ensure execution success.
Our Programs - Deep Dive
Every program at Recursion is a product of our Recursion OS. Our wholly-owned programs are built on unique biological insights surfaced through the Recursion OS and target diseases where: i) the disease-biology is well defined and ii) there is high unmet medical need, there are no approved therapies, or there are significant limitations to existing treatments. Several of our programs target indications with market opportunities expected to be near to or in excess of $1.0 billion in annual sales and we are preparing four programs to enter Phase 2 or Phase 2/3 clinical trials within the first three quarters of 2022 and a fourth program to enter a Phase 1 clinical trial within the second half of 2022.
7 Haas JV, Eastwood BJ, Iversen PW, et al. Minimum Significant Ratio – A Statistic to Assess Assay Variability. 2013 Nov 1 [Updated 2017 Nov 20]. In: Markossian S, Sittampalam GS, Grossman A, et al., editors. Assay Guidance Manual [Internet]. Bethesda (MD): Eli Lilly & Company and the National Center for Advancing Translational Sciences; 2004.
45
Table of Contents
Figure 35. Examples of current Recursion programs falling into our First, Second and Next Generation paradigms. The earliest iterations of the Recursion OS leveraged brute-force search (where small molecules were tested directly in the context of each disease model we built) and used a small molecule library restricted primarily to known chemical entities. Programs arising from this iteration of the Recursion OS are deemed First Generation Programs. As we developed our chemistry capabilities and new chemical entity library at Recursion, Second Generation Programs arose, though the throughput needed to screen large libraries of new chemical entities presents a powerful but relatively inefficient solution. Today, most of our new programs, as well as new partnerships or expansions of prior partnerships, are Next Generation Programs, whereby we use our maps of biology to navigate to novel or unexpected relationships between molecules (known or new chemical entities) and then validate those predictions in our wet-labs.
•Recursion’s First Generation of Potential Medicines. The following programs represent the novel use of a known chemical entity discovered using early iterations of the Recursion OS.
◦REC-994 for the treatment of cerebral cavernous malformation, or CCM— Phase 2a enrolling patients at the time of filing. Orphan Drug Designation granted in the US and EU.
◦REC-2282 for the treatment of neurofibromatosis type 2, or NF2—expected Phase 2/3 initiation in Q2 2022. Orphan Drug Designation in the US and EU, as well as Fast-Track Designation in the US, have been granted.
◦REC-4881 for the treatment of familial adenomatous polyposis, or FAP—expected Phase 2 initiation in Q3 2022. Orphan Drug Designation granted in the US.
◦REC-3599 for the treatment of GM2 gangliosidosis, or GM2—expected Phase 2 initiation in 2024.
•Recursion’s Second Generation of Potential Medicines. The following programs arose from a brute-force approach leveraging either an expanded internal new chemical entity library or a partner new chemical entity library.
46
Table of Contents
◦REC-3964 for the treatment of C. difficile colitis— expected Phase 1 initiation in 2H, 2022
◦REC-64917 for Neural or Systemic Inflammation
◦Multiple simultaneous programs in fibrosis advancing with Bayer
•Recursion’s Next Generation of Potential Medicines. The following programs represent a promising subset of known or new chemical entities discovered and developed using the latest Recursion OS mapping and navigating tools.
◦REC-65029 and derivatives or functionally related series for the Treatment of HRD-negative Ovarian Cancer by leveraging a potentially novel target insight
◦REC-648918 and derivatives or functionally related series to enhance anti-tumor immune response leveraging a potentially novel target insight (Target Alpha)
◦REC-2029 for the treatment of Wnt-mutant Hepatocellular carcinoma
◦REC-14221 and derivatives or functionally related series for the treatment of solid and hematological malignancies using indirect MYC inhibition
◦REC-64151 and derivatives or functionally related series for the treatment of immune checkpoint resistance in KRAS/STK11 mutant non-small cell lung cancer
◦Potential future programs in fibrosis with Bayer or in neuroscience or a single oncology indication with Roche and Genentech
In addition to the programs highlighted above, we are actively developing dozens of additional programs which may prove to be drivers of our future growth. As we have significantly expanded our chemistry capabilities in the last year, and continue to invest deeply in these key elements of the Recursion OS, moving forward we expect that the vast majority of our new programs will be part of our Next Generation of potential programs discovered using our tools for mapping and navigating biology. We believe that the number of potential programs we can generate with our Recursion OS is key to the future of our company, as a greater volume of validated programs has a higher likelihood of creating value. The speed at which our OS generates a large number of product candidates is important, since traditional drug development often takes a decade or more. In addition, we believe that our large number of potential programs makes us an attractive partner for larger pharmaceutical companies. The static or declining level of R&D output at many large companies means that they have an ongoing need for new projects to fill their pipelines.
Figure 36. The power of our Recursion OS as exemplified by the breadth of active research and development programs. We have an expansive pipeline of internally-developed programs spanning multiple therapeutic areas and consisting of both new uses for existing compounds and new chemical entities, or NCEs, under active research and development. All populations are US and EU5 incidence unless otherwise noted. EU5 is
47
Table of Contents
defined as France, Germany, Italy, Spain and the UK. (1) Prevalence for hereditary and sporadic symptomatic population. (2) Annual US and EU5 incidence for all NF2-driven meningiomas. (3) Worldwide prevalence; conducting dose optimization study in animal model with a potential trial start in 2024 (4) US and EU5 prevalence (5) Our program has the potential to address a number of indications with systemic or neural inflammatory components. We have not finalized a target product profile for a specific indication. (6) Our program has the potential to address a number of indications driven by MYC alterations, totaling 54,000 patients in the US and EU5 annually. We have not finalized a target product profile for a specific indication. (7) Our program has the potential to address a number of indications in this space.
First Generation Program - REC-994 for Cerebral Cavernous Malformation
Early Discovery Late Discovery Preclinical Phase 1 Phase 2 Phase 3
Summary
REC-994 is an orally bioavailable, superoxide, scavenger small molecule being developed for the treatment of CCM. In Phase 1 SAD and MAD trials in healthy volunteers directed and executed by us, REC-994 demonstrated excellent tolerability and suitability for chronic dosing. CCM is among the largest rare disease opportunities and has no approved therapies. We recently enrolled the first patient in a Phase 2 double-blind, placebo-controlled, safety, tolerability and exploratory efficacy study.
Disease Overview
CCM is a disease of the neurovasculature for which approximately 360,000 patients in the US and EU5 have been diagnosed or suffer symptoms. Less than 30% of patients with CCM experience symptoms, resulting in the disease being severely underdiagnosed and suggesting that well more than 1 million patients may have the disease in the US and EU5. CCM and its hallmark vascular malformations are caused by inherited or somatic mutations in any of three genes involved in endothelial function: CCM1, CCM2, or CCM3. Approximately 20% of patients have a familial form of CCM that is inherited in an autosomal dominant pattern. Sporadic disease in the remaining population is caused by somatic mutations that arise in the same genes. CCM manifests as vascular malformations of the spinal cord and brain characterized by abnormally enlarged capillary cavities without intervening brain parenchyma. Patients with CCM lesions are at substantial risk for seizures, headaches, progressive neurological deficits and potentially fatal hemorrhagic stroke. Current non-pharmacologic treatments include microsurgical resection and stereotactic radiosurgery. Given the invasive and risky nature of these interventions, these options are reserved for a subset of patients with significant symptomatology and/or easily accessible lesions. Rebleeds and other negative sequelae of treatment further limit the effectiveness of these interventions. There is no approved pharmacological treatment that affects the rate of growth of CCM lesions or their propensity to bleed or otherwise induce symptoms. CCM can be a severe disease resulting in progressive neurologic impairment and a high risk of death due to hemorrhagic stroke.
Product Concept
We are developing an orally bioavailable small molecule therapeutic designed to alleviate neurological symptoms associated with CCM and potentially reduce the accumulation of new lesions. REC-994 is an orally bioavailable small molecule superoxide scavenger with pharmacokinetics supporting once-daily dosing in humans. Mechanistically, the reduction of endothelial superoxide species has been shown to reverse the cellular pathogenesis of the disease. In addition, REC-994 exhibits anti-inflammatory properties which could be beneficial in reducing disease-associated pathology. Preclinical data have demonstrated benefit on acute to subacute disease-relevant hemodynamic parameters such as vascular permeability and vascular dynamics. Chronic administration in rodent genetic models of CCM has demonstrated long-term benefit in reduction of lesion number and/or size. REC-994 was well tolerated at up to 800 mg daily dosing in healthy human subjects enrolled in our Phase 1 study, and there were no severe adverse events at any dose tested. The safety results of the Phase 1 studies we executed support continued evaluation of REC-994 in a Phase 2 study. We licensed global rights for the data underlying our novel usage of REC-994 from the University of Utah in February 2016 and have obtained orphan drug designation in the US and EU.
Preclinical
48
Table of Contents
The novel use of REC-994 for CCM was discovered leveraging knock-down of the disease gene CCM2 in primary human endothelial cells using the earliest form of the Recursion OS. In secondary orthogonal assays, REC-994 reversed defects in human endothelial cell-cell junctional integrity, a functional phenotype associated with the loss of CCM2.
REC-994 was subsequently tested in two endothelial-specific knockout mouse models for the two most prevalent genetic causes, Ccm1 and Ccm2. These mouse models faithfully recapitulate the CNS cavernous malformations of the human disease. Mice treated with REC-994 demonstrated a decrease in lesion number and/or size compared to vehicle treated controls. Notably, 24-hour circulating plasma levels of REC-994 in this in vivo experiment were consistent with exposures seen in humans at a 200 mg daily dose.
Figure 37. REC-994 reduces lesion severity in chronic mouse models of CCM Disease. Mice treated with REC-994 demonstrated a statistically significant decrease in the number of small-size lesions, with a trend toward a decrease in the number of mid-size lesions.
Clinical
We recently enrolled the first patient in a Phase 2 double-blind, placebo-controlled, safety, tolerability and exploratory efficacy study.
We conducted a SAD study in 32 healthy human volunteers using active pharmaceutical ingredients with no excipients in a Powder-in-Bottle dosage form. Results showed that systemic exposure (Cmax and AUC) generally increased in proportion to REC-994 after both single and multiple doses. Median Tmax and t1/2 appeared to be independent of dose. There were no deaths or SAEs reported during this study and no TEAEs that led to the withdrawal of subjects from the study. These data supported a MAD study in healthy human volunteers.
The MAD study was conducted in 52 healthy human volunteers and was designed to investigate the safety, tolerability, and PK of multiple oral doses of REC-994, to bridge from the Powder-in-Bottle dosage form to a tablet dosage form, as well as to assess the effect of food on PK following a single oral dose. Overall, multiple oral doses of REC-994 were well tolerated in healthy male and female subjects at each dose level administered in this study. There appeared to be no dose-related trends in TEAEs, vital signs, ECGs, pulse oximetry, physical examination findings, or neurological examination findings. Pharmacokinetic results support once-daily oral dosing with the tablet formulation.
49
Table of Contents
Table 3. Summary Statistics for Plasma REC-994 Pharmacokinetic Parameters – Overall MAD Cohorts.
Table 4. Summary of Adverse Events from Phase 1 Multiple Ascending Dose Study. AE=adverse event; MAD=multiple ascending dose; SAE=serious adverse event; TEAE=treatment-emergent adverse event
We recently enrolled the first patient in an exploratory Phase 2 double-blind placebo-controlled, safety, efficacy and pharmacokinetics study of REC-994 in the treatment of symptomatic CCM. The study is enrolling patients with symptomatic CCM at least 18 years of age with anatomic CCM lesions demonstrated by MRI. The primary objective of the Phase 2 study will be to assess the safety and tolerability of daily dosing of a low and high dose group of REC-994 over 12 months, compared to placebo, in patients with symptomatic CCM. Exploratory secondary endpoints will include assessment of patient reported outcomes, imaging assessments, as well as established composite scales for neurological signs and symptoms.
Currently, there is no development or regulatory precedent or pathway for CCM drug development. We will undertake an exploratory Phase 2 to inform a pivotal trial design with guidance from the FDA.
50
Table of Contents
Figure 38. Phase 2 clinical trial schematic for REC-994. Planned Phase 2 trial design to assess the efficacy and safety of REC-994 in patients with symptomatic CCM.
Competitors
There are two investigator-initiated clinical studies underway to study marketed therapeutics in CCM patients.
•Investigators at the University of Chicago are evaluating the efficacy of atorvastatin, or Lipitor, on reduction in hemorrhage rate in patients with CCM.
•Investigators at the Mario Negri Institute for Pharmacological Research in Italy are evaluating the efficacy of the approved beta blocker propranolol in reducing lesions and clinical events.
To our knowledge, the REC-994 program is the only industry-sponsored therapeutic program in clinical trials for CCM. If approved, REC-994 would be the first pharmacologic disease-modifying treatment for CCM, one of the largest areas of unmet need in the rare disease space.
First-Generation Program - REC-2282 for Neurofibromatosis Type 2
Early Discovery Late Discovery Preclinical Phase 1 Phase 2 Phase 3
Summary
REC-2282 is a small molecule HDAC inhibitor being developed for the treatment of NF2-mutant meningiomas. The molecule has been well tolerated, including in patients dosed for multiple years, and potentially reduced cardiac toxicity that differentiates it from other HDAC inhibitors. In contrast to approved HDAC inhibitors, REC-2282 is both CNS-penetrant and orally bioavailable. We expect to enroll the first patient in an adaptive, parallel group, Phase 2/3, randomized, multicenter study in the second quarter of 2022.
Disease Overview
Neurofibromatosis Type 2 (NF2) is an autosomal dominant, inherited, rare, tumor syndrome caused by loss-of-function mutations in the NF2 tumor suppressor gene, which encodes the cell signaling regulator protein merlin. Loss of NF2 results in growth of the hallmark tumors that characterize this disease: vestibular schwannomas (VS) and meningiomas. The tumor types of VS and meningiomas seen in NF2 are among the most common in neuro-oncology. In addition, NF2 mutations give rise to spontaneous meningiomas, mesotheliomas, and underlie subsets of additional tumor types. Combined, we believe NF2-driven meningiomas occur in approximately 33,000 patients per year in the US and EU5. Patients with NF2 are diagnosed typically in their late teens or early 20s and present with hearing loss which is usually unilateral at the time of onset, focal neurological deficits, and symptoms relating to increasing intracranial pressure. Although the course of disease progression is highly variable, most patients are rendered deaf, and many will eventually need wheelchair assistance due to progressive neurological decline. The standard of care is surgery or radiosurgery and patients may require multiple operative procedures during their lifetime. Although surgery or radiation can be effective in controlling tumor growth, most surgical procedures result in morbidity related to neurological deficits based on the location of the tumor. Hearing loss, facial nerve palsy, and moderate facial nerve dysfunction are also common surgical outcomes. Radiation can induce malignant transformation which in turn makes surgery more complex. In addition, tumors may recur post-surgical resection along with the growth of new tumors. NF2-associated tumors and treatment related morbidity can lead to earlier than expected mortality. If left untreated, NF2-driven tumors can result in death resulting from rising intracranial pressure.
Product Concept
51
Table of Contents
REC-2282, is an orally bioavailable, CNS-penetrating, pan-histone deacetylase, or HDAC, inhibitor with PI3K/AKT/mTOR pathway modulatory activity. By comparison to marketed HDAC inhibitors, REC-2282 is uniquely suited for patients with NF2, and NF2-mutant CNS tumors, due to its oral bioavailability and CNS-exposure. NF2 disease is driven by mutations in the NF2 gene, which encodes an important cell signaling modulator, merlin. Loss of merlin results in activation of multiple signaling pathways converging on PI3K/AKT/mTOR among others. Human clinical pharmacodynamic data supports the role of REC-2282 in inhibiting activity of multiple aberrant signaling pathways in NF2-deficient tumors. HDAC inhibitors induce growth arrest, differentiation, and apoptosis of cancer cells. We obtained a global license for REC-2282 from the Ohio State Innovation Foundation in December 2018. Orphan drug designation for REC-2282 in NF2 has been granted in the US and EU. Fast Track Designation for REC-2282 in NF2-mutated meningioma was granted in the US in 2021.
Figure 39. REC-2282 acts on an important pathway in tumor development to inhibit the growth of tumor cells. A potential mechanism of action of REC-2282 in NF28.
Preclinical
The novel use of REC-2282 for NF2was discovered leveraging the knock-down of the disease gene NF2 in human cells in the Recursion OS. We did not see similarly robust responses in the context of many other tumor suppressor genes studied, suggesting some specificity of the mechanism in the context of NF2 loss of function.
Figure 40. Impact of REC-2282 on NF2 model in the Recursion OS. REC-2282 reversed the effects of knock-down of NF2 in primary human cells using our phenomics assay.
8Adapted from Petrilli and Fernández-Valle. Role of Merlin/NF2 inactivation in tumor biology. Oncogene 2016 35(5):537-48
52
Table of Contents
After we discovered the novel use of REC-2282 for NF2 using our platform, we performed a literature search to better understand the molecule and validate disease models. At that time, we discovered that REC-2282 had been shown to inhibit in vitro proliferation of vestibular schwannoma, or VS, and meningioma cells by inducing cell cycle arrest and apoptosis at doses that correlate with AKT inactivation. In preclinical models, REC-2282 inhibited the growth of primary human VS and NF2-deficient mouse schwannoma cells, as well as primary patient-derived meningioma cells and the benign meningioma cell line, Ben-Men-1.
In animal models of NF2, REC-2282 suppressed in vivo growth of an NF2-deficient mouse vestibular schwannoma allograft. In addition, REC-2282 suppressed in vivo growth human vestibular schwannoma xenograft models in mice fed, either a standard diet of rodent chow, or chow formulated to deliver 25 mg/kg/day REC-2282 for 45 days. REC-2282 also suppressed the growth of an orthotopic mouse model of NF2-deficient meningioma using luciferase-expressing Ben-Men-1 meningioma cells. These animal data served as a functional and orthogonal validation of our platform findings.
Figure 41. REC-2282 prevents tumor growth in Vestibular Schwannoma xenografts. REC-2282 significantly reduces the mean size of VS xenografts in SCID-ICR mice. Error bars shown are the 95% CI. P=0.006. Adapted from Jacob, 2011. REC-2282 also suppressed the growth of Ben-Men-1-LucB tumor xenografts as measured by tumor luminescence. Adapted from Burns, 2013.
Clinical
We expect to initiate a parallel group, two-staged, Phase 2/3, randomized, multicenter clinical trial within the next quarter.
Previous clinical work conducted in investigator-initiated trials and trials sponsored by Arno Therapeutics (no longer a licensor of Ohio State University for this molecule) includes human exposure to REC-2282, previously referred to as AR-42. A total of three completed studies in adult human subjects were conducted in the United States in patients with solid or hematological malignancies. Published data from Ohio State University reflects that a total of 77 patients were treated with REC-2282 in doses ranging from 20 mg to 80 mg three times a week for three weeks followed by one week off-treatment in four-week cycles. Multiple patients were treated for multiple years using this dosing regimen at the 60 mg dose and the longest recorded treatment duration is 4.4 years at the 40 mg dose. The majority of adverse events were transient cytopenia that did not result in dose reduction or stoppage. The MTD in patients with solid tumors was determined to be 60 mg. The REC-2282 plasma exposure in patients with hematological malignancies and solid tumors generally increased with increasing doses. There were no consistent signs of plasma REC-2282 accumulation across a 19-day administration period nor obvious differences in PK between hematologic and solid tumor patients.
In another early Phase 1 pharmacodynamic study by Ohio State University, it appears that REC-2282 suppressed aberrant activation of ERK, AKT, and S6 pathways in vestibular schwannomas resected from treated NF2 patients. These results may be difficult to achieve with single pathway inhibitors of ALK or MEK.
We are planning to initiate an adaptive, parallel group, two-staged, Phase 2/3, randomized, multicenter study to evaluate the efficacy and safety of REC-2282 in patients with progressive NF2 mutated meningiomas with underlying NF2 disease and sporadic meningiomas with documented NF2 mutations.
53
Table of Contents
The study is designed to accelerate the path to potential product registration by allowing for initiation of a confirmatory Phase 3 study prior to full completion of Phase 2. It is a combined Phase 2-3 study design, beginning with a Proof-of-Concept Phase 2 portion in which 20 adult subjects (and up to nine adolescent subjects) will begin treatment on two active dose arms. Subject safety will be monitored by an independent Data Monitoring Committee, which will apply dose modifications and stopping rules as indicated. After all 20 adult subjects have completed six months of treatment, an interim analysis will be performed for the purposes of 1) determination of go/no-go for Phase 3 portion of the study, 2) selection of the dose(s) to carry forward, 3) re-estimation of sample size for the planned Phase 3, and 4) agreement from FDA to initiate Phase 3. Subjects in the Phase 2 will continue treatment for up to 26 months total and then have the option to enroll in an Open Label Extension study. The Phase 3 portion of the design currently requires recruitment of an additional 60 subjects (adult and potentially adolescent subjects), who will receive treatment for up to 26 months. The planned primary endpoint is Progression-Free Survival (PFS).
Figure 42. Phase 2/3 clinical trial schematic for REC-2282. Planned Phase 2/3 trial design to assess the efficacy and safety of REC-2282 in meningioma patients.
Competitors
There are currently fouractive programs in clinical development targeting NF2-driven brain tumors.
•Brigatinib, an approved ALK inhibitor for NSCLC from Takeda Pharmaceuticals, is in Phase 2 for NF2 disease meningioma, vestibular schwannoma and ependymoma.
•Crizotinib, an ALK/ROS1 inhibitor, is being studied in an investigator sponsored Phase 2 study in progressive vestibular schwannoma in NF2 patients.
•Selumetinib, a MEK inhibitor from AstraZeneca, is being studied in a Phase 2 trial for NF2 related tumors.
•GSK2256098, a FAK inhibitor from GlaxoSmithKline, is being studied in a basket Phase 2 for meningiomas with a variety of targeted therapies and genetic alterations, including NF2 mutation.
First Generation Program - REC-4881 for Familial Adenomatous Polyposis (FAP)
Early Discovery Late Discovery Preclinical Phase 1 Phase 2 Phase 3
Summary
REC-4881 is an orally bioavailable, non-ATP-competitive allosteric small molecule inhibitor of MEK1 and MEK2 being developed to reduce polyp burden and progression to adenocarcinoma in FAP patients. REC-4881 has been well tolerated in prior clinical studies, consistent with the intended use and has a gut-localized PK-profile in humans that is highly advantageous for FAP, and potentially other APC-driven gastrointestinal tumors. We expect to enroll the first patient in a Phase 2, double-blind, randomized, placebo-controlled basket trial in the third quarter of 2022.
54
Table of Contents
Disease Overview
FAP is a rare tumor syndrome affecting approximately 50,000 patients in the US and EU5 with no approved therapies. FAP is caused by autosomal dominant inactivating mutations in the tumor suppressor gene APC, which encodes a negative regulator of the Wnt signaling pathway. FAP patients develop polyps and adenomas in the colon, rectum, rectal pouch, stomach, and duodenum throughout life. These growths have a high risk of malignant transformation and can give rise to invasive cancers of the colon, stomach, duodenum, and rectal tissues. Standard of care for patients with FAP is colectomy in late teenage years. Without surgical intervention, affected patients will progress to colorectal cancer by early adulthood. Post-colectomy, patients receive endoscopic surveillance every 6-12 months to monitor disease progression given the ongoing risk of malignant transformation.
Despite surgical management, the need for effective pharmacological therapies for FAP remains high due to continued risk of duodenal and desmoid tumors post-surgery. These tumors occur in the majority of patients and surgical resection of these tumors can be associated with significant morbidity. NSAIDs, such as sulindac or celecoxib, are sometimes used to treat these tumors, but have limited efficacy and do not impact pre-cancerous lesions. While surgical management and surveillance have improved the prognosis for FAP patients, desmoid tumors remain a major cause of death in patients with FAP following colectomy.
Product Concept
Our REC-4881 candidate is an orally bioavailable, non-ATP-competitive allosteric small molecule inhibitor of MEK1 and MEK2 (IC50 2-3 nM and 3-5 nM, respectively) that has demonstrated potent reduction in polyps and dysplastic adenomas, in the Apcmin mouse model of FAP. In a previous Phase 1 clinical study run by Millennium Pharmaceuticals, 51 patients with solid tumors were treated with REC-4881 and did not demonstrate the typical ocular toxicities associated with this class. REC-4881 exhibits extremely low hepatic metabolism and its primary route of elimination is through biliary excretion and gastrointestinal elimination, which may allow it to achieve preferential exposure at tumor sites in the duodenum and lower gastrointestinal tract with reduced systemic exposures and toxicity. We obtained a global license for REC-4881 from Takeda Pharmaceuticals in May 2020. Orphan drug designation for REC-4881 in FAP and APC-driven tumors was granted by the FDA in 2021.
Preclinical
The novel use of REC-4881 for FAP was discovered leveraging knock-down of the FAP disease gene APC in human cells using the Recursion OS. We validated our findings using tumor cell lines and spheroids grown from human epithelial tumor cells with a mutation in APC. REC-4881 inhibited both the growth and organization of spheroids in these models and, in tumor cell lines, had well over a 1,000-fold selectivity range in cells harboring APC mutations.
Figure 43. Impact of REC-4881 on anAPC model on the Recursion OS. REC-4881 reversed the effects of knockdown of APC in primary human cells using our phenomics assay.
We subsequently evaluated REC-4881 in a disease relevant preclinical model of FAP. Mice harboring truncated Apc, or Apcmin, were treated with multiple oral daily doses of REC-4881 or celecoxib (as a comparator) over an eight-week period. Mice treated with celecoxib had approximately 30% fewer polyps than did those treated with vehicle, whereas mice treated with 1 mg/kg or 3 mg/kg REC-4881 exhibited approximately 50% fewer polyps than vehicle-treated mice. Mice that were treated with 10 mg/kg REC-4881, the highest dose tested, exhibited an approximately 70% reduction in total polyps.
55
Table of Contents
Figure 44. REC-4881 reduces GI polyp count in the Apcminmousemodel of FAP. GI polyp count after oral administration of indicated dose of REC-4881, celecoxib, or vehicle control for 8 weeks. Polyp count at start of dosing reflects animals sacrificed at the start of study (15 weeks of age). P < 0.001 for all REC-4881 treatment groups versus vehicle control.
In FAP, polyps arising from mutations in APC may progress to high-grade adenomas through accumulation of additional mutations and eventually to malignant cancers. To evaluate the activity of REC-4881 on both benign polyps and advanced adenomas, gastrointestinal tissues from mice treated with REC-4881 were histologically evaluated and polyps were classified as either benign or high-grade adenomas. While celecoxib reduced the growth of benign polyps in the model, a large proportion of polyps that remained were dysplastic. By contrast, treatment with REC-4881 specifically reduced not only benign polyps, but also precancerous high-grade adenomas, a finding with the potential for translational significance.
Figure 45. Disease progression of FAP begins with mutations in APC. Progression of benign APC-mutant polyps to high-grade adenomas and eventually malignant tumors occurs following the accumulation of additional genetic alterations9.
9 Adapted from http://syscol-project.eu/about-syscol/
56
Table of Contents
Figure 46.REC-4881 reduces high-grade adenomas in the Apcmin mouse model of FAP9. Quantification of high-grade adenomas versus total polyps based on blinded histological review by a pathologist. While celecoxib reduces benign polyps, the majority of remaining lesions are high grade adenomas. By contrast, REC-4881 reduces both polyps and high-grade adenomas.
REC-4881 is a non-ATP-competitive and specific allosteric small molecule inhibitor of MEK1 and MEK2. Studies have shown that mitogen-activated protein kinase signaling, or MEK, and extracellular signal-regulated kinase, or ERK signaling is activated in adenoma epithelial cells and tumor stromal cells, including fibroblasts and vascular endothelial cells.
In addition, genomic events resulting in alteration of mitogen-activated protein kinase signaling, or MAPK, such as activating mutations in KRAS, are frequent somatic events that promote the growth of adenomas in FAP. Therefore, suppression of aberrant MAPK signaling in adenomas of FAP with REC-4881 has the potential to regress or slow the growth of these tumors by acting on core pathways driving their growth.
Clinical
Millennium Pharmaceuticals previously conducted clinical work including human exposure using REC-4881, then referred to as TAK-733. A total of 51 patients were included in the Phase 1 study, which demonstrated that REC-4881 had a manageable toxicity profile up to the maximum tolerated dose, or MTD, of 16 mg dosed on days one to 21 of 28-day treatment cycles. The most common adverse events were dermatitis acneiform rash (53%), fatigue (36%), and diarrhea (31%), consistent with other MEK inhibitors. No dose-limiting toxicities, or DLTs, were observed in patients who received REC-4881 in doses from 0.2 mg to 8.4 mg. Four patients experienced DLTs of grade 3 dermatitis acneiform at doses of 12 mg (n=1), 16 mg (n=1), and 22 mg (n=2). Importantly, REC-4881 demonstrated fewer adverse ocular side effects compared to approved drugs in this class. Our preclinical data in FAP support a low dose cohort in the Phase 2 trial in the dosing range where DLTs were not experienced in the prior Phase 1 (0.2 - 8.4 mg).
Study C20001 was a Phase 1, multicenter, open-label, dose-escalation, first-in-human clinical trial designed to evaluate the safety, pharmacokinetics, and pharmacodynamics of TAK-733 (now REC-4881) in patients with advanced, nonhematologic malignancies and melanoma.Abbreviations: AUC0-24hr: area under the plasma concentration versus time curve from zero to 24 hours; CLss/F: apparent oral clearance; Cmax: maximum plasma concentration; CV%: percent coefficient of variation; NC: not calculable; Std Dev: standard deviation; Tmax: time of first observed maximum concentration. Mean and geometric mean were calculated if 2 or more individual parameter values. Median, Std Dev, CV%, min, and max were calculated if 3 or more individual parameter values. Summary statistics for PK parameters are not presented in this table for 0.2, 0.4, 0.8, and 1.6 mg cohorts, as N<3 in these cohorts. The number of patients (n) may differ from the total N in each dose cohort depending on the parameter and day. Source: CSR C20001
We plan to initiate a Phase 2, randomized, double-blind, placebo-controlled basket trial to evaluate efficacy, safety and pharmacokinetics of REC-4881 in classical FAP patients. We expect to initiate this Phase 2 clinical trial by the end of Q3 2022.
•The study will be conducted in classical FAP who are at or over 18 years of age at the time of enrollment.
57
Table of Contents
•The study will be conducted in two parts. Part A will evaluate the effects of food and dosing interval on the pharmacokinetics of REC-4881 (as the drug has not been studied in patients with colectomy previously). Part 2 will evaluate the efficacy, safety and pharmacokinetics of REC-4881.
•Patients from three subpopulations will be randomized into two active and one placebo group and treated for 12 months.
•The study will assess tumor response endpoints in patients treated with REC-4881 versus placebo.
Figure 47. Clinical trial schematic for REC-4881. Planned Phase 2 clinical trial to assess the efficacy, safety and pharmacokinetics of REC-4881 in patients with classical FAP.
Key Competitors
There are four primary therapeutic approaches in clinical development for FAP; all are focused on reduction in colorectal polyposis.
•Guselkumab (Tremfya) is an IL-23 human monoclonal antibody, or mAb, in Phase 2 development by Janssen Pharmaceuticals which is hypothesized to reduce cytokine production, inflammation, and tumor polyp development.
•Eicosapentaenoic acid-free fatty acid is a polyunsaturated fatty acid currently in Phase 3 development for FAP by S.L.A. Pharma AG. Eicosapentaenoic acid-free fatty acid is hypothesized to reduce polyp formation due to its activity as a competitive inhibitor of arachidonic acid oxidation.
•A combination of Eflornithine and sulindac (CPP1X/Sulindac) is in development by Cancer Prevention Pharma for FAP and, in a recent Phase 3 study, the incidence of disease progression with the combination was not significantly lower than either drug alone. The company submitted an NDA in June 2020, and it remains under review. The company withdrew their MAA application in October 2021.
•Encapsulated rapamycin, or eRAPA, is currently in Phase 2 development by Emtora Biosciences for FAP and is hypothesized to reduce tumor formation through its inhibitory effect on the mTOR pathway.
First Generation Program - REC-3599 for GM2 Gangliosidosis
Early Discovery Late Discovery Preclinical Phase 1 Phase 2 Phase 3
58
Table of Contents
Summary
REC-3599 is an orally bioavailable, selective, potent small molecule inhibitor of Protein Kinase C beta, or PKCß, and Glycogen synthase kinase 3 beta, or GSK3ß being developed for the treatment of infantile GM2 gangliosidosis. REC-3599 has demonstrated strong reduction of pathogenic biomarkers GM2 and lipofuscin levels in cells derived from patients with multiple different mutations in either HEXA or HEXB, referred to as Tay-Sachs or Sandhoff Disease, respectively. We are planning to generate additional pharmacodynamic and efficacy data in a sheep model of GM2. We anticipate enrolling the first patients in a Phase 2 trial in infantile GM2 patients in 2024.
Disease Overview
GM2 gangliosidosis, or GM2, is a lysosomal storage disease affecting approximately 400 patients in the US and EU5. The disease is caused by mutations in either HEXA or HEXB genes which encode subunits of the lysosomal beta-hexosaminidase enzyme. Mutations in HEXA lead to Tay-Sachs disease and mutation in HEXB lead to Sandhoff disease. GM2 presents during infancy, childhood, or later in life depending upon the degree of genetic deficiency and is classified by the period of onset: Infantile onset, Juvenile onset, and Late-onset Tay-Sachs or Sandhoff Disease. Patients with infantile GM2 are diagnosed in the first year of life and exhibit rapidly progressing neurological decline, associated with neuronal lysosomal dysfunction and GM2 accumulation, resulting in complete neurological disability and premature death in the first few years of life. Some of the earliest observed signs include retinal abnormalities and exaggerated startle reflex within the first six-months after birth. Affected infants may achieve some motor milestones at close to expected normal developmental age up to about 12-months; however, they will ultimately lose any gained motor skills, including basic skills such as the ability to turn over, sit, crawl, and swallow, by the age of 18-24 months and usually succumb to their disease prior to age four. There are no approved symptomatic or disease modifying treatments for the disease. Standard of care for these patients is supportive interventions, including seizure control with anticonvulsants, assisted feeding through a nasogastric tube, or percutaneous endoscopic gastrostomy, and, ultimately, ventilatory support. While progression of the disease remains rapid, supportive care can provide some improvement in the survival for patients with infantile GM2.
Product Concept
We are developing a small molecule therapeutic as monotherapy or in combination with gene therapy to slow progression of neurological decline in patients with infantile GM2. REC-3599 is an orally bioavailable, CNS-penetrant small molecule inhibitor of PKCß with additional inhibitory activity on GSK3ß. In preclinical studies, REC-3599 demonstrated potent and concentration-dependent reduction of GM2-ganglioside accumulation and sphingolipid-associated autofluorescence in infantile GM2 patient-derived fibroblast models at IC50s suitable for human dosing. REC-3599 is hypothesized to play a dual role in modulating lysosomal biogenesis through inhibition of GSK3ß while also stimulating cellular autophagy through inhibition of PKCß. Eli Lilly previously studied REC-3599, then referred to as ruboxistaurin, in adult patients with diabetic retinopathy, including in Phase 3 clinical trials. The compound has been dosed in over >2,500 adult human subjects with treatment durations as long as two years. REC-3599 has been well tolerated in adult human subjects, supporting its evaluation in this rare and devastating infantile neurological disease. We executed a relevant in vivo pharmacodynamic study and juvenile rodent toxicology studies at the request of the FDA to help bridge entry into pediatric populations.
In 2015, Eli Lilly out-licensed the rights for ruboxistaurin to Chromaderm; we subsequently licensed the global rights to ruboxistaurin from Chromaderm for all systemic uses in December 2019. We obtained pediatric rare disease designation for REC-3599 in GM2 in 2020. We plan to seek orphan drug designation for REC-3599 in GM2 following the generation of additional pharmacodynamic data in a HEXA deficient GM2 animal model and completion of the planned Phase 2 study in patients with infantile GM2 gangliosidosis.
Preclinical
The novel use of REC-3599 for GM2 was discovered leveraging the Recursion OS using a knockdown of the GM2 disease gene HEXB in human cells.
59
Table of Contents
Figure 48. Impact of REC-3599 on HEXB model. REC-3599 reversed the effects of knockout of HEXB in human cells using our phenomics assay.
In Tay-Sachs and Sandhoff diseases, the loss of function of ß-hexosaminidase results in the accumulation of GM2 gangliosides and lipofuscin in the lysosome. Exposure of infantile GM2 patient fibroblast lines to REC-3599 resulted in a reduction in GM2 ganglioside aggregates, total GM2 levels, and lipofuscin-associated autofluorescence to levels comparable to apparently healthy control-derived fibroblast lines. These data are consistent with an improvement in lysosomal function resulting from REC-3599 exposure.
REC-3599 was initially developed as an inhibitor of PKCß; however, the compound also demonstrates weaker but significant inhibitory activity against GSK3ß. GSK3ß is a known inhibitor of lysosomal biogenesis, and inhibition of GSK3ß has been shown to lead to increased lysosomal production and function by activating transcription of lysosomal genes regulated by transcription factor TFEB. Additionally, inhibition of GSK3ß leads to pro-survival autophagic signaling through TFEB. In parallel, results support the role of PKCß as an inhibitor of cellular autophagy, a key cellular process in lysosomal-mediated degradation that is impaired in lysosomal storage diseases. Thus, the dual action of REC-3599 in modulating lysosomal biogenesis through inhibition of GSK3ß while also stimulating cellular autophagy through inhibition of PKCß, may underlie the unique activity of REC-3599 in human cellular models of GM2.
Figure 49. Infantile patient cells show reduced disease-specific activity when treated with increasing doses of REC-3599. Infantile Tay-Sachs and Sandhoff disease patient fibroblasts exhibit higher: mean GM2 fluorescence (left panel), aggregate counts (middle panel), and autofluorescent substrate accumulation (right panel).
Clinical
We are planning to initiate a Phase 2 clinical trial in Infantile GM2 in 2024.
Previous clinical work conducted by Eli Lilly includes considerable human exposure to REC-3599, including a total of 37 studies in adult human subjects with a total of 4,094 participants: 26 clinical pharmacology studies (including a QT study) that included a total of 573 adult subjects that have established the absorption, distribution, metabolism, excretion, pharmacodynamics, and tolerability of REC-3599; and 11 placebo-controlled studies that included a total of 3,521 adult subjects with diabetes and moderate to severe non-proliferative retinopathy. An additional 2 randomized, placebo-controlled trials in adults with diabetic macular edema and 1 safety and PK study in patients with diabetes were conducted after the initial marketing applications and included an additional 1,069 adult subjects.
60
Table of Contents
In the clinical pharmacology studies, single doses of REC-3599 up to 256 mg and multiple daily doses up to 128 mg given over two weeks were taken by healthy subjects. In double-blind, placebo-controlled, safety and efficacy studies, REC-3599 was administered at daily doses of 4, 8, 16, and 32 mg for ≥ 36 months, and 64 mg for ≥ 12 months. In the Eli Lilly clinical trials REC-3599 has been well tolerated at the doses administered to adults.
Safety information provided in Eli Lilly’s NDA 22005 supports the safety profile of REC-3599 in adult patients. The summary of safety conclusions was as follows: Most adverse events were noted to be mild to moderate severity and did not lead to discontinuation of study drug; the safety profile of REC-3599 was similar regardless of age, gender, ethnicity, and type of diabetes. REC-3599 32 mg administered once per day was the intended dose for patients with diabetic retinopathy. In Eli Lilly’s clinical program, the incidence of patients with at least 1 serious adverse event, or SAE, was lower in 32 mg REC-3599 treated patients compared with placebo; the pattern of SAEs did not suggest any organ-specific or systemic toxicity. Analyses of laboratory measures, vital signs, and ophthalmic safety assessments revealed no clinically significant safety concerns.
Upon satisfactory completion of in vivo pharmacodynamic and efficacy studies in the sheep HEXA model, we expect to initiate an open-label Phase 2a study evaluating the efficacy, safety, tolerability, pharmacokinetics, and pharmacodynamics of REC-3599 in patients with Infantile GM2 gangliosidosis. We expect to initiate a Phase 2a clinical trial in 2024.
•The study will be conducted in pediatric patients with confirmed diagnosis of infantile GM2.
•The study will consist of four periods: screening, dose escalation, treatment, and follow-up. The anticipated treatment period is 36 months.
•An interim analysis is planned after 12 months of treatment of the last enrolled patient.
•We will track the achievement of development milestones, neurological function, and quality of life using established and validated composite scales.
Figure 50. Phase 2 clinical trial schematic for REC-3599. Planned Phase 2a clinical trial to assess the efficacy, safety, and pharmacokinetics of REC-3599 in patients with Infantile GM2.
Key Competitors
Key competitors to the REC-3599 program consist of two therapeutic categories, gene therapies and small molecule substrate reduction therapies. Two companies are developing AAV-based gene therapies to restore functional beta-hexosaminidase enzymes by gene delivery:
•Taysha Gene Therapies is developing an AAV-based gene therapy, TSHA-101. The program is currently in Phase 2.
•Sio Gene Therapies is also developing an AAV-based gene therapy, AXO-AAV-GM1/GM2. The program is currently in Phase 1/2.
Two companies are developing small molecule substrate reduction therapies:
61
Table of Contents
•Sanofi is developing Venglustat as an orally bioavailable small molecule hypothesized to reduce substrate accumulation in GM2 and other lysosomal storage diseases. The program is currently in Phase 3 studies in patients with late-onset GM2.
•IntraBio is developing N-Acetyl-L-Leucine as an orally bioavailable amino-acid ester. The program has completed a Phase 2 study.