Kosmos today made commercially available an operational intelligence platform that makes use of machine learning algorithms to discover the root cause of an outage or issues that are adversely affecting an IT environment.

Fresh off raising $5 million in seed funding, Kosmos CEO Sanjay Gidwani said that the Kosmos Operational Intelligence platform will apply artificial intelligence (AI) within a correlation engine to generate deterministic output using signals collected from platforms such as Jira, Salesforce, GitHub, ServiceNow, Datadog, Grafana and Splunk.

The overall goal is to reduce the amount of time required to investigate IT incidents from hours to a few minutes, he added.

Most IT teams, when investigating incidents, typically set up a “war room” from which participants are excused when they are able to prove that the application or IT infrastructure they are responsible for is not part of the problem. That process, however, can require days to complete.

The reason for that is that information about applications, code changes, and service incidents remains fragmented across myriad tools that do not share context. Instead, the Kosmos Operational Intelligence platform automatically surfaces correlations that, once confirmed by humans, are used to create a continuous intelligence loop that learns more about the IT environment with each incident investigated.

In effect, Kosmos is making a case for an operational intelligence platform that employs AI to significantly reduce mean time to resolution (MTTR). That capability can be especially critical when there is an outage that is having a direct impact on the organization’s ability to generate revenue.

With the rise of AI, correlating the relationships between elements of complex IT environments has now become a solvable problem, said Gidwani.

It’s not clear how much time is being wasted by IT teams in war rooms in a given year, but the need to correlate issues faster is becoming more pressing now that organizations depend on IT to function. The challenge is that IT environments are now too large for most humans to truly understand and correlate all the dependencies that exist without help from AI. While it may be possible to rely on probabilistic generative AI tools to identify some of those relationships, the Kosmos platform relies on machine learning algorithms to generate more reliable deterministic outputs.

Each IT organization will need to determine for itself to what degree it needs to rely on AI to correlate relationships in a way that reduces MTTR, but, in general, there are more important things most IT teams could be doing if they were not spending time in a war room trying to essentially prove their innocence.

Regardless of the approach to IT incident management, the primary goal will always remain trying to prevent issues from arising in the first place. However, the chances that IT environments are not going to experience some sort of outage are near zero, so the challenge and the opportunity now is how best to leverage advances in AI to reduce the total amount of downtime an organization might experience simply because some dependency between applications and IT infrastructure wasn’t fully understood or appreciated.