Abstract
Enterprise identity governance is designed around a limited number of central Identity Providers (IdPs), such as Okta, Microsoft Entra ID and Ping ID. Enterprise identity governance is based on as few as necessary identity providers (IdPs) that federate human sign-on using single sign-on (SSO), SAML and OpenID Connect (OIDC). However, a significant and ever-increasing portion of authentication is not sent across this plane. All of these credentials sit outside of the central IdP's view and control: local operating-system accounts, statically embedded service-account credentials, unmanaged software-as-a-service (SaaS) logins, developer-issued API keys, OAuth grants negotiated between objects directly in the running applications, and ever more independent artificial-intelligence (AI) agents that log on via local or scripted access paths. This group of ungoverned, un-Federated authentication occurs, in this paper, in a manner that is referred to as “shadow identity.” We formalize the problem of detecting shadows, place it in the context of three related fields in the exterior community: Non-human identity (NHI) governance, workload attestation standards, and graph-based identity analytics, propose a layered detection architecture, the Multi-Source Shadow Identity Detection (MSSID) framework. MSSID combines network- or endpoint-level telemetry, cross-source log correlation, construction of heterogeneous identity graphs, and three complementary detection models: a coverage-gap model, a graph-neural-network-style (GNN-style) behavioral anomaly model and an attestation-consistency model, to uncover authentication activity that has escaped the radar of the centralized governance approach. In addition to the proposal for the architecture, we deploy three reference copies of all three proposed detection models and test them end-to-end in a controlled synthetic test domain with 1,590 identity nodes and 190 shadow instances, distributed across all six taxonomy classes we define, which is currently not available in the public domain. The operating system used in this system has implemented it with a precision of 0.89, a recall of 0.72, an F1 score of 0.79, a cover ratio of 0.90 compared with the seeded ground truth, a mean simulated time to discovery of 1.0 day for the fusion operating point, and 90.5% accuracy for attribution. The results per class support the main design claims of the framework: The predominant model for none of the six classes of shadow identity; and the contributions of the three models are remarkably complementary rather than redundant. These results are not reported as a measure of validated results in the field; rather, the reports are the only measure of internal coherence and potential for production deployment at testbed scale. Finally, we conclude with a discussion about evasion resistance, the implications for governance, limitations, and a research agenda for detecting shadows of identities under attestation in a privacy-preserving and explainable way.