The Global Phenotype and the Shape of Mycelial Possibility
The Shape of Mycelial Possibility
When phenotype is discussed in practice, it is most often treated as a single or few features that can be cleanly described: a morphology, a growth rate, a texture, a yield, a mechanical response. This shorthand is useful, but it is also incomplete, insomuch as any single cultivation instance of a fungus is not the complete phenotype of the organism. It is an instantaneous expression, a single coordinate, within a much larger space of possibility.
It is critical to recognize that phenotyping is a deep and rich technical space that can vary meaningfully in its scope and definition depending on context and goals, so I will recognize that, from here forward, I am discussing the concept from a personally relevant stance. As such, at its most fundamental level, the global phenotype is the full range of physical and functional expressions that a fungal genotype can produce across the complete dimensionality of its operating context. It is not what the fungus does under a particular set of conditions, but what it can do as environment, growth matrix, spatial context, geometry, history, scale, and time are varied. Any individual morphology, material property, or performance metric is simply a local sample of this broader space.
Biology has long recognized that phenotype is conditional, where reaction norms describe how traits shift across environments (Schlichting & Pigliucci, 1998; Via & Lande, 1985; West-Eberhard, 2003). Developmental landscapes describe stability, canalization, and path dependence (Ferrell, 2012; Huang, 2012). Adaptive and performance landscapes describe how form maps to function or fitness (Arnold, 2003). High-dimensional phenotyping and phenomics seek to measure phenotype as a composite object rather than a single trait (Furbank & Tester, 2011; Houle et al., 2010). The global phenotype, as I am using the term here, sits comfortably within this intellectual terrain. I am using it as an operational concept: a generalized, multivariate, high-dimensional reaction norm for describing how a strain expresses physical and functional possibility across an operator-defined parameter space. It is a way of holding these ideas together while keeping the organism’s whole-body physical expression at the center of attention. In this framing, phenotype is not a scalar or even a vector, but a high-dimensional response surface. That surface contains all expressible physical states, not just the stable or commonly observed ones: planes, basins, ridges, and edge states. Individual physical features are not independent levers; they are correlated outputs that emerge together as the organism interprets its conditions. Each feature is a dependent variable, and each cultivation samples the surface differently.
Therefore, the global phenotype comprises n emergent physical manifolds (structured surfaces of correlated physical expression) as a function of k operational dimensions defined by the operator.
Crucially, the global phenotype is the map, a response surface. Any discrete physicality observed in a single run is just one coordinate on the surface. Treating that coordinate as representative of the strain as a whole collapses a landscape into a point (the definition of missing the forest for the tree). Through this lens, what looks like stability may simply be residence in a broad basin, whereas unpredictability may be traversal across a steep gradient. What looks like a breakthrough may be a narrow ridge that is difficult to re-enter. The surface itself has structure. It contains regions of robustness where behavior is repeatable across a wide range of conditions. It also contains fragile pockets where extreme or unusual phenotypes appear. It contains history-dependent pathways, where outcome depends not only on present conditions but on the trajectory taken to arrive there. And it contains scale-dependent distortions, where the same nominal parameters map to different outcomes as volume, geometry, or spatial gradients change.
Seen this way, the global phenotype defines a physically embodied opportunity space. It describes the envelope of behaviors a strain can express and, by extension, the envelope of addressable physicality for human use. The central challenge of mycelium engineering is not only identifying a single “best” phenotype for the particular application, but understanding which regions of the global phenotype are useful, accessible, and robust under real-world variability. In occupying this space, the mycelium is actively interpreting inputs. The global phenotype encodes how the organism tends to respond, how broad or narrow its response ranges are, and how it negotiates tradeoffs across space and time. In this sense, it is also a map of fungal agency. As such, phenotyping becomes a practice of learning this agency in order to find overlap between fungal responsiveness and human intent; a sort of negotiation, looking for the win-win.
Phenotype as a Composite Variable
If the global phenotype describes the full landscape of possible expression, then phenotype itself must be treated, for the purposes of this essay and this kind of applied strain characterization, as a composite variable (Houle et al., 2010; Kitano, 2002). Any attempt to reduce phenotype to one or two traits (growth rate, density, strength, yield) may be adequate depending on operational context, but inevitably it risks collapsing correlated behavior into an incomplete summary. And here the danger lies in missing risks or opportunities connected to correlated responses. For instance, could selecting for a particular desired quality inadvertently also select for a quality that makes scaling more difficult? Or can the same quality be achieved at multiple positions in the defined operational space, where one position co-expresses with another more advantageous trait?
In this framing, phenotype is inferred through the joint behavior of many physical features expressed together under specific conditions (i.e. physical manifolds). Each feature we measure is a projection of a deeper, coupled response.
Formally, the global phenotype can be described as a set of n emergent response manifolds, each expressed as a function of k operator-defined dimensions:
Global Phenotype = {Rᵢ(O) | i = 1, 2, …, n; O ∈ ℝᵏ}
where O represents the operational space defined by the operator, including environmental parameters, nutritional variables, geometry, scale, time, and perturbations. Each Rᵢ represents an emergent physical manifold: a structured surface of correlated physical responsiveness expressed by the organism.
The operational dimensions (k) are largely chosen, bounded, and named by the experimenter or process designer respective to practical operability through the experimental or scaled operational space; critically, this is defined by operational reality. The physical response dimensions (n) are not. They emerge from how the organism integrates those inputs, and their dimensionality is typically unknown. What we call “phenotyping” is the act of sampling and approximating these emergent manifolds through measurement, as such selection of physical qualities for measurement is critical for ensuring that the relevant physical responsiveness to the operational and product context is sampled; the actual physical responsiveness is not defined, but the degree of visibility of the organisms responsiveness is defined by the operator.
Each manifold represents a coordinated pattern of behavior: morphology, bioefficiency, or mechanics that move together as conditions vary. These features are correlated because they are jointly produced by the same underlying biological interpretation of context. In practice, this means that individual features should rarely be treated as independent levers. They are better understood as coupled readouts of position within a response surface. The language of manifolds is useful here precisely because it avoids implying independence. A manifold is a structured surface embedded within a higher-dimensional space. When we measure multiple physical features, we are not uncovering separate dimensions so much as sampling points on one or more manifolds that constrain how those features can vary together.
At the same time, the operational dimensions themselves are not guaranteed to be orthogonal. For example, environmental conditions (temperature, humidity, oxygen availability) and growth geometry routinely interact. Treating them as independent inputs may not accurately represent biological truth, whereas the composite-variable framing allows this reality to be acknowledged without demanding perfect experimental isolation.
From a practical standpoint, this shifts the goal of phenotyping. The objective is no longer to measure “the phenotype,” but to (1) identify which features are meaningfully correlated, (2) understand how those correlations shift across operational dimensions, and (3) determine which regions of the resulting response surfaces are broad, stable, and useful.
This is why, when it comes to practically understanding the breadth of a strain’s capacity, feature engineering, dimensionality reduction, and clustering are the tools by which composite phenotypes become legible and allow complex, coupled physical behavior to be compared across strains, conditions, and scales. In this sense, treating phenotype as a composite variable is a practical acknowledgment of how fungal systems actually behave, and how they must be engaged if strain characterization and improvement is to be organized around both opportunity and robustness.
Practical Dimensioning of the Global Phenotype
The first step in working with the global phenotype is deciding how it will be sampled. Because the global phenotype cannot be observed directly, it must be approximated through a deliberate combination of physical measurements and controlled variation in operating conditions, consistent with the broader logic of high-dimensional phenotyping and phenomics (Furbank & Tester, 2011; Houle et al., 2010). This is the act of defining the coordinate system in which the organism’s behavior will be observed.
The operator begins by selecting a set of physical features that meaningfully describe the form and function of interest. These features should be chosen, respectively, according to both what is practically achievable and what gives the widest potential view of physicality relevant to the application. This might include, as an example mycelial density, biological efficiency, total mycelium mass, compressive strength, elastic modulus, or tensile strength. This could then represent a reasonably practical sampling regime that would provide sensitivity, intuitively, to multiple distinct physical qualities (gross physicality, metabolic efficiency, aggregate mechanical behaviors).
Each of these measurements represents one axis of the physical response space, but importantly, are not necessarily assumed to be orthogonal. Changes in density will likely correlate with mechanical performance; biological efficiency may correlate with total biomass (or may not); stiffness and strength may covary. At this stage, the goal is not to resolve causality or isolate effects, but simply to ensure that the chosen features collectively span the physical behaviors the operator cares about.
Next, the operator defines the operational dimensions across which those physical features will be sampled. These are the parameters the operator can set, vary, or bound respective to real world targets, conditions, and limitations. Continuing the same example, this operational space might include temperature, substrate composition (potentially itself multi-dimensional: carbon to nitrogen ratios, particle size, nutrient density), ambient humidity, and growth time.
At this point it is worth acknowledging a meaningful departure between an academic research mentality and a private, process- and product-driven R&D mentality. Academic training often emphasizes exhaustive characterization of a knowledge space; the aspiration to fully resolve a system before drawing conclusions. Taken literally, this implies that complete factorial exploration of a parameter space is the ideal. From this vantage point, efficient sampling strategies can appear secondary, tools of convenience rather than core competencies. In practice, this framing can underemphasize statistical learning and design-of-experiments methods that are explicitly designed to avoid full factorial sampling. If the goal is total resolution of a human-agnostic problem space, sparse designs look like compromises. Why invest in efficient sampling if the ideal outcome is exhaustive enumeration?
Private, product- and process-driven R&D operates under a different constraint set. Here, full factorial sampling is almost never appropriate, because the objective is almost never exhaustive understanding. The objective is timely, defensible decision-making under uncertainty, in service of real operational targets; the shape of the solution space, not every solution. This drives a different competency profile, prioritizing sparse, information-dense sampling, rapid elimination of irrelevant regions of parameter space, and learning strategies that maximize insight per unit effort rather than completeness. Seen through this lens, defining the operator-driven parameter space is not an attempt to fully characterize the organism’s human-agnostic physical potential, but rather an intentional act of alignment; to efficiently map fungal responsiveness to contexts where that physicality can be converted into value. This is why dimension reduction, adaptive sampling, and human-driven parameterization are the mechanisms by which biological possibility becomes legible and actionable under real-world constraints.
Sampling the Global Phenotype
The full Cartesian product of the operational space is rarely practical. And again, in service to maximizing actionable learning power per unit effort, the operational space must be sampled strategically. Several families of experimental design are appropriate here:
Orthogonal designs emphasize independence between factors and are part of the broader design-of-experiments toolkit (Montgomery, 2019). Space-filling designs, such as Latin hypercube sampling or Sobol’ sequences, aim to cover the operational space without privileging any single axis (McKay et al., 1979; Sobol’, 1967). Adaptive or sequential designs use early results to guide subsequent experiments toward regions of interest, either to maximize learning or to target specific performance envelopes (Chaloner & Verdinelli, 1995).
The choice among these approaches reflects the practical real-world interest of the operator, the organization, and the technology; broad exploration of phenotypic possibility, focused interrogation of a suspected region, or efficient discrimination between strains. The resulting dataset does not yet reflect insight, but reflects raw physical-operational structure that can be projected to insight. Each cultivation produces a position in the k-dimensional operational space, and a corresponding observation in the n-dimensional physical response space. Taken together, these runs form a sparse but intentional sampling of the global phenotype.
Feature Engineering, Dimension Reduction, and Projection
Once the global phenotype has been sampled through a structured combination of physical measurements and operational variation, the next task is to make that structure legible. At this point the dataset may exist as a table: rows of cultivation runs, columns of operational parameters and physical responses. The global phenotype is structurally present in this table, but only implicitly, and extracting its shape requires transforming raw measurements into representations that reflect coordinated biological behavior rather than isolated variables.
This begins with feature engineering. The physical response space may not necessarily be most informative in its raw form. Individual measurements (density, strength, modulus, efficiency) are often noisy, scale-dependent, and partially redundant. Feature engineering is the process of constructing or transforming variables so that they better represent the structure relevant to downstream modeling or interpretation (Kuhn & Johnson, 2019). In this context, the goal is to construct derived variables that better reflect meaningful physical behavior. This may involve normalization, ratios, composite metrics, or aggregation across spatial or temporal windows. The intent is to express existing biological coordination more clearly, and reflects a significant opportunity for the operator to drive relevance and focus of the downstream understanding of the global phenotype. Similarly, it is often useful to partition the response space into conceptually related subsets. Mechanical features may be considered together, metabolic features together, morphological features together. This partitioning is a pragmatic step that allows different aspects of physicality to be examined without immediately collapsing everything into a single representation.
With features defined, the analysis turns to dimension reduction. The physical response space is n-dimensional, but the effective dimensionality of meaningful variation is often much lower. Many features move together because they are jointly constrained by underlying biological processes. Dimension reduction methods aim to identify those dominant directions of coordinated variation. Principal component analysis (PCA) provides a useful starting point. PCA identifies orthogonal directions in feature space that capture the greatest variance in the data, making it a common tool for reducing dimensionality while preserving interpretable structure (Jolliffe & Cadima, 2016). When applied to physical response features, these components often correspond to interpretable axes of behavior: overall densification, stiffness-strength tradeoffs, growth-efficiency regimes. Importantly, PCA provides a compact, linear projection that makes correlation structure, and ultimately phenotypic structure, visible.
Multiple correspondence analysis (MCA) is appropriate when physical features are categorical or discretized rather than continuous (Greenacre & Blasius, 2006), while nonlinear approaches such as autoencoders can be useful when response manifolds are strongly curved or when interactions between features cannot be well approximated by linear projections (Hinton & Salakhutdinov, 2006). Because these methods often sacrifice direct interpretability for expressive power, it becomes important to actively recover meaning from their outputs. This can be supported by examining feature loadings or square cosine (cos²) contributions to understand which original variables are most strongly represented in each reduced dimension, by computing cross-correlations between original physical features and the resulting latent axes, and by probing how small perturbations in input features propagate through the reduced representation. Used this way, dimension reduction does not obscure biological meaning but reframes it, allowing complex, coupled physical behavior to be expressed in a compact form while retaining a clear connection to measurable properties.
Dimension reduction can also be applied to the operational space itself. In complex substrate formulations or multi-factor environmental regimes, operational parameters may be correlated or redundant. Reducing the effective dimensionality of the operational space can clarify which combinations of conditions are functionally distinct from the organism’s perspective.
Once reduced representations are available, the global phenotype can be examined through projection. Reduced physical response coordinates are mapped back onto the operational space, allowing the operator to see how coordinated physical behavior varies as conditions change. This is where the abstract notion of response surfaces becomes concrete. Through these projections, the qualitative shape of the global phenotype begins to emerge.
Broad plateaus appear where physical behavior is relatively insensitive to changes in operational parameters. Basins form where multiple perturbations converge toward similar outcomes, indicating regions of stability. Ridges appear where small changes in conditions produce sharp shifts in physical response, often associated with novel performance but low robustness. Boundaries and folds reveal history-dependent pathways, where access to a region depends on how it is approached rather than on final conditions alone. These features are practical descriptors of how physical responsiveness is distributed across the operational space. A flat region suggests tolerancing and manufacturability, a narrow ridge suggests opportunity coupled with risk, a fragmented surface suggests sensitivity to uncontrolled variation. Importantly, these qualities belong to the global phenotype of a specific strain under a specific framing, not to the organism in isolation; fungal agency and physical expression mapped to human interest and opportunity.
At this stage, the analysis remains strain-specific. The goal is intra-strain from an operational framing: identifying which regions of the global phenotype are broad or narrow, stable or fragile, accessible or elusive, and ultimately manufacturable and useful. Only once this internal structure is visible does it become meaningful to compare strains as differently shaped phenotypic landscapes occupying the same operational frame.
Comparing Strains Through Phenotypic Shape
Once the internal structure of a strain’s global phenotype has been made visible (the basins, plateaus, ridges, and boundaries) it becomes possible to compare strains in a way that moves beyond point metrics. At this point, comparison is how different strains occupy, traverse, and stabilize within the same operational frame. The key is to place multiple strains into a shared representational space, using the same feature engineering pipeline, dimensionality reduction, and operational framing. Each strain’s response manifolds can be projected into a common coordinate system. This alignment is critical, without it differences in apparent performance may simply reflect differences in measurement, scaling, or framing rather than genuine phenotypic divergence.
Once projected into a shared space, differences between strains begin to express themselves as differences in phenotypic shape, not just magnitude. The global phenotype can now be treated as a geometric object, allowing spatial language to be used directly in comparison. Some strains occupy broad plateaus of acceptable performance, corresponding to large contiguous regions of phenotypic space which signal robustness across wide ranges of operational conditions. Others express narrow ridges of high performance coupled to steep gradients and sharp drop-offs, signaling sensitivity to perturbation and limited tolerance to drift. Still others occupy largely disjoint regions of phenotypic space, achieving similar endpoint metrics through different routes across the operational landscape. In this framing, strains can be compared by the extent, continuity, and geometry of their phenotypic regions; by the area or volume of space they occupy, by the degree and location of overlap between their useful regions, or by distance between dominant response manifolds, rather than by isolated point measurements.
Cluster analysis becomes useful at this point as a tool for identifying regions of phenotypic similarity and divergence within this shared geometry. At its core, cluster analysis groups observations based on proximity or similarity within a defined feature space, allowing structure to emerge without prespecifying categories (Jain, 2010). Clustering can be applied to reduced physical response coordinates, to operational projections, or to combined representations that encode both response and context. In each case, clusters do not represent fixed “types” of strains so much as shared regimes of behavior: zones of the global phenotype where strains behave similarly despite differences elsewhere. Strains that cluster together may share robustness properties, scaling characteristics, or failure modes, even if their peak metrics differ under any single condition. Used this way, clustering helps distinguish between strains whose phenotypic shapes meaningfully overlap in target regions and those whose apparent similarity masks different operational realities.
This approach also makes offsets explicit. Two strains may achieve comparable values for a target property, such as strength or yield, but do so in different regions of the operational space. The ability to detect and describe offsets between strains respective to a common target can suggest how forgiving a process will be, how sensitive it is to drift, and how easily it can be transferred across scales or facilities. Importantly, shape-based comparison reframes what it means for one strain to be “better” than another. Superiority is no longer defined by peak performance alone, but by the size, accessibility, and stability of useful regions within the global phenotype. A strain with a slightly lower maximum performance but a broad, flat operating region may be far more valuable than one that occasionally reaches extreme values along a narrow and fragile ridge. As scale increases and operational eccentricities become harder to avoid, operators can explicitly weigh absolute point performance against the robustness of the surrounding phenotypic region. Under these conditions, a strain that expresses slightly lower peak performance but maintains that performance across a wider and more forgiving region of the global phenotype may ultimately outperform a strain whose higher peak is accessible only under tightly constrained conditions.
This framing also allows tradeoffs to be seen clearly. Improvements in one region of the global phenotype may compress or distort another. A strain optimized for rapid densification may sacrifice mechanical tunability. A strain with highly predictable morphology may lose access to extreme architectures. By comparing shapes rather than points, these tradeoffs become visible rather than latent. At this point, comparison becomes an exercise in alignment between biological behavior and human intent. Different applications privilege different regions of the global phenotype: robustness for manufacturing, sensitivity for exploration, tunability for design optionality. The same strain may be well suited to one context and poorly suited to another because its phenotypic shape does not align with the target region of interest.
In this sense, comparative phenotyping completes the loop. It closes the distance between biological possibility and operational decision-making by allowing strains to be evaluated as dynamical systems embedded within an operational landscape. What emerges is a map that makes it possible to choose strains with intention, to anticipate failure modes, and to design processes that work with the organism’s intrinsic structure rather than against it.
Strain Development as Agency Mapping
Strain development can be practiced as the optimization of single-point phenotypes, focusing on narrow or isolated features measured under a nominal or “standard” condition. But to view it this way is, in my opinion, somewhat unfortunate. It misses much of the point, and much of the opportunity, of working with fungi as a bioprocess tool. It collapses a rich, expressive physical system into a handful of outcomes, often sacrificing physical novelty and silently passing risk downstream to stages where, once detected, it becomes significantly more disruptive (and expensive) to manage.
This is why viewing strain development as a practice of mapping fungal agency can be exciting, derisking, and powerful. It reframes the work around learning how a strain tends to interpret conditions: how broadly or narrowly it responds, where it is stable, where it is sensitive, and what tradeoffs it expresses as it moves through operational space. The global phenotype becomes the working map of that agency, while the operator-defined parameter space becomes the practical frame within which that agency can be sampled, explored, and ultimately made actionable. In applied mycelium engineering, we are rarely trying to resolve the fungus in a human-agnostic sense. We are trying to align biological possibility with human boundary conditions so that the organism’s physical responsiveness can be converted into reliable value.
How do we design with, and maximize value from, the full vocabulary of fungal physicality? The answer is literacy; learning the organism’s expressive range well enough to collaborate with it. The global phenotype offers a language for treating fungal behavior as a structured opportunity space rather than a collection of surprises, and strain development becomes the disciplined act of shaping that space so that fungal agency and human intent converge.
References
Arnold, S. J. (2003). Performance surfaces and adaptive landscapes. Integrative and Comparative Biology, 43(3), 367–375. https://doi.org/10.1093/icb/43.3.367
Chaloner, K., & Verdinelli, I. (1995). Bayesian experimental design: A review. Statistical Science, 10(3), 273–304. https://doi.org/10.1214/ss/1177009939
Ferrell, J. E., Jr. (2012). Bistability, bifurcations, and Waddington’s epigenetic landscape. Current Biology, 22(11), R458–R466. https://doi.org/10.1016/j.cub.2012.03.045
Furbank, R. T., & Tester, M. (2011). Phenomics: Technologies to relieve the phenotyping bottleneck. Trends in Plant Science, 16(12), 635–644. https://doi.org/10.1016/j.tplants.2011.09.005
Greenacre, M., & Blasius, J. (Eds.). (2006). Multiple correspondence analysis and related methods. Chapman and Hall/CRC. https://doi.org/10.1201/9781420011319
Hinton, G. E., & Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks. Science, 313(5786), 504–507. https://doi.org/10.1126/science.1127647
Houle, D., Govindaraju, D. R., & Omholt, S. (2010). Phenomics: The next challenge. Nature Reviews Genetics, 11(12), 855–866. https://doi.org/10.1038/nrg2897
Huang, S. (2012). The molecular and mathematical basis of Waddington’s epigenetic landscape: A framework for post-Darwinian biology? BioEssays, 34(2), 149–157. https://doi.org/10.1002/bies.201100031
Jain, A. K. (2010). Data clustering: 50 years beyond K-means. Pattern Recognition Letters, 31(8), 651–666. https://doi.org/10.1016/j.patrec.2009.09.011
Jolliffe, I. T., & Cadima, J. (2016). Principal component analysis: A review and recent developments. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 374(2065), Article 20150202. https://doi.org/10.1098/rsta.2015.0202
Kitano, H. (2002). Systems biology: A brief overview. Science, 295(5560), 1662–1664. https://doi.org/10.1126/science.1069492
Kuhn, M., & Johnson, K. (2019). Feature engineering and selection: A practical approach for predictive models. Chapman and Hall/CRC. https://doi.org/10.1201/9781315108230
McKay, M. D., Beckman, R. J., & Conover, W. J. (1979). A comparison of three methods for selecting values of input variables in the analysis of output from a computer code. Technometrics, 21(2), 239–245. https://doi.org/10.1080/00401706.1979.10489755
Montgomery, D. C. (2019). Design and analysis of experiments (10th ed.). Wiley.
Schlichting, C. D., & Pigliucci, M. (1998). Phenotypic evolution: A reaction norm perspective. Sinauer Associates.
Sobol’, I. M. (1967). On the distribution of points in a cube and the approximate evaluation of integrals. USSR Computational Mathematics and Mathematical Physics, 7(4), 86–112. https://doi.org/10.1016/0041-5553(67)90144-9
Via, S., & Lande, R. (1985). Genotype-environment interaction and the evolution of phenotypic plasticity. Evolution, 39(3), 505–522. https://doi.org/10.1111/j.1558-5646.1985.tb00391.x
West-Eberhard, M. J. (2003). Developmental plasticity and evolution. Oxford University Press. https://doi.org/10.1093/oso/9780195122343.001.0001