How Should a CTO Manage a Crisis? [+3 Case Studies][2026]
In an era where technology underpins every critical business function, the ability of a Chief Technology Officer to manage a crisis can determine whether an organization survives or collapses under pressure. From ransomware attacks and social engineering schemes to source code theft and identity system breaches, the threats CTOs face today are more sophisticated and damaging than ever before. This article explores how CTOs should approach crisis management through 10 actionable strategies covering incident response, transparent communication, data-driven decision making, and continuous improvement. To bring these strategies to life, the article also examines three real-world case studies involving MGM Resorts, Okta, and Slack, drawing lessons from some of the most consequential technology crises in recent years. DigitalDefynd has compiled these insights to equip technology leaders with the knowledge and frameworks needed to protect their organizations and lead with confidence when a crisis strikes.
Table of Contents
Real-World 3 Case Studies
- MGM Resorts: Social engineering ransomware attack causing $100 million loss in 2023
- Okta: Customer support system breach exposing session tokens of 18,000+ clients in 2023
- Slack: Employee token theft exposing private GitHub code repositories in late 2022
How Should a CTO Manage a Crisis?
1. Foster a culture of continuous improvement
- Learning from mistakes
- Innovation in response
- Collaborative improvement efforts
2. Implement robust incident response plans
- Preparation is key
- Team readiness
- Tools and technologies
3. Transparent and timely communication
- Internal coordination
- Stakeholder communication
- Post-crisis review
4. Utilize data-driven decision making
- Data analysis for insights
- Predictive analytics for proactive management
5. Establish clear leadership and responsibility
- Define roles and responsibilities
- Leadership visibility
6. Collaborate across departments and with external experts
- Cross-departmental collaboration
- Engagement with external experts
7. Develop and maintain a business continuity plan
- Comprehensive coverage
- Regular updates and testing
8. Prioritize employee training and awareness
- Ongoing training programs
- Crisis simulation exercises
9. Strengthen security and risk management
- Regular security audits
- Risk management framework
10. Leverage technology for real-time monitoring and alerts
- Implementation of monitoring tools
- Alert systems
How Should a CTO Manage a Crisis? [3 Case Studies]
1. MGM Resorts: Social Engineering Ransomware Attack Causing $100 Million Loss in 2023
Challenge
In September 2023, MGM Resorts International fell victim to a devastating ransomware attack orchestrated by two cybercriminal groups, Scattered Spider and ALPHV/BlackCat. The attackers used social engineering tactics, impersonating an MGM employee on a 10-minute call to the IT help desk to obtain administrator credentials. This single breach gave them access to MGM’s Okta and Microsoft Azure environments, from which they deployed ransomware across more than 100 ESXi hypervisors. The attack crippled operations across more than 30 properties, taking slot machines, digital room keys, online reservation systems, and payment platforms offline for approximately 10 days. The incident exposed a critical gap in MGM’s human-centered security protocols, underscoring how even a technically advanced organization can be undone by social engineering when employee awareness and help desk verification procedures are insufficient.
Solution
a. Immediate System Isolation: Upon detecting unusual network activity on September 10, 2023, MGM’s technology leadership moved swiftly to shut down critical systems to contain the breach. While disruptive, this decision successfully prevented threat actors from accessing customer bank account numbers or payment card information, limiting the scope of data compromise.
b. Incident Response Mobilization: MGM engaged external cybersecurity experts and law enforcement agencies, including the FBI, to investigate the breach, conduct forensic analysis, and guide the recovery process. This cross-departmental and multi-agency collaboration was central to containing further damage.
c. Refusal to Pay Ransom: Unlike its rival Caesars Entertainment, which reportedly paid a ransom during a simultaneous attack, MGM’s leadership decided against paying. This decision aligned with guidance from cybersecurity experts and government agencies, prioritizing long-term security integrity over short-term resolution.
d. Post-Crisis Technology Overhaul: MGM committed $50 million to upgrade endpoint protection, cloud security infrastructure, and employee training programs. The technology leadership also enforced stricter identity verification protocols at the IT help desk to close the social engineering vulnerability that enabled the attack.
Result
The financial and reputational toll on MGM was severe. The company reported approximately $100 million in losses to its third-quarter 2023 earnings, with an additional $10 million in one-time expenses covering legal fees, technology consulting, and third-party advisory services. MGM’s stock dropped 4.1% over two trading days following the breach. By September 20, 2023, full system restoration was confirmed. The incident became a defining case study in how CTOs must integrate human-layer security, not just technical defenses, into comprehensive crisis management frameworks.
Related: Types of Chief Technology Officers
2. Okta: Customer Support System Breach Exposing Session Tokens of 18,000+ Clients in 2023
Challenge
In October 2023, Okta, a leading identity and access management provider serving more than 18,000 customers globally, disclosed a significant breach of its customer support case management system. Attackers gained access using a stolen credential and extracted HTTP Archive (HAR) files that customers had uploaded during support sessions. These files contained active session tokens and cookies, which the threat actors used to infiltrate customer Okta tenants without needing passwords or bypassing multi-factor authentication. The initial compromise occurred around September 28, 2023, but Okta’s security team did not detect it until October 13, giving attackers approximately two weeks of undetected access. The breach rippled outward, impacting high-profile customers including 1Password, Cloudflare, and BeyondTrust. For a company whose core business is securing digital identities, the incident posed severe reputational and operational risks, exposing a fundamental contradiction between Okta’s security promise and its own internal vulnerabilities.
Solution
a. Credential Invalidation and Containment: Upon confirming the breach on October 13, Okta’s technology leadership immediately invalidated all compromised session tokens and reset affected credentials to cut off attacker access. This rapid containment step was critical to preventing further lateral movement across customer environments.
b. Customer Notification and Transparency: Okta publicly disclosed the incident on October 19, 2023, notifying approximately 134 directly affected customers. The security team provided detailed guidance on reviewing Okta System Logs for unusual administrator activity, particularly access from unexpected locations or unmanaged devices.
c. HAR File Policy Overhaul: Recognizing that the support workflow itself was the attack vector, Okta restructured its internal processes for handling sensitive files uploaded by customers. New procedures required sanitizing HAR files of session tokens before they were shared with or stored in the support system.
d. Long-Term Security Enhancements: Following the breach, Okta implemented enhanced monitoring across its support systems, enforced stricter access controls on customer data, and strengthened its incident communication protocols. Multi-factor authentication requirements and privileged access management controls were tightened across internal systems to prevent credential-based intrusions.
Result
The breach impacted 134 confirmed customer organizations and triggered an immediate decline in Okta’s stock value. The delayed detection window of approximately two weeks drew sharp criticism from the cybersecurity community and affected customers, particularly BeyondTrust, which had alerted Okta to the suspicious activity on October 2 but received no acknowledgment for 17 days. The incident reinforced for CTOs everywhere that even security-first organizations must apply the same rigorous monitoring standards internally that they prescribe to their customers, and that delayed crisis communication compounds reputational damage far beyond the technical breach itself.
Related: How Can CTOs Achieve Work-Life Balance?
3. Slack: Employee Token Theft Exposing Private GitHub Code Repositories in Late 2022
Challenge
On December 31, 2022, Slack, a Salesforce-owned business communication platform used by an estimated 18 million users worldwide, disclosed a security incident affecting its externally hosted GitHub code repositories. The breach, which occurred on December 27, involved threat actors stealing a limited number of Slack employee tokens and using them to gain unauthorized access to private repositories. The attackers downloaded private code from these repositories before the intrusion was detected. While Slack confirmed that neither its primary codebase nor any customer data was present in the compromised repositories, the theft of proprietary code raised legitimate concerns about what vulnerabilities or configuration details the stolen files might reveal to future attackers. The incident also drew scrutiny over Slack’s disclosure timing, as the breach occurred on December 27 but was only publicly announced on December 31, and the security notice was initially marked to prevent search engine indexing, raising transparency concerns.
Solution
a. Immediate Token Invalidation: Upon being notified of suspicious activity on December 29, 2022, Slack’s technology leadership moved quickly to invalidate all stolen employee tokens, cutting off the threat actor’s access pathway to the GitHub repositories and preventing further data extraction.
b. Credential Rotation Across Systems: As a precautionary measure, Slack rotated all relevant secrets and credentials connected to the compromised tokens. This step ensured that even if the attackers attempted to leverage related access points, those pathways were no longer viable.
c. Forensic Investigation and Scope Assessment: Slack engaged in a thorough investigation to determine the full scope of the breach, confirming that the downloaded repositories contained no customer data, no means to access customer data, and no components of Slack’s primary codebase. This scoping was critical to accurately communicating risk levels to stakeholders.
d. Strengthened Token and Access Management: Following the incident, Slack reinforced its internal processes around API key security, employee token management, and access controls for externally hosted repositories. Guidance emphasized the use of secure environment variables and secret management services over embedded credentials.
Result
Slack confirmed that no customer data was exposed and that its core platform and services remained fully operational throughout the incident. However, the breach highlighted a growing industry-wide pattern of attackers targeting employee credentials and tokens as entry points rather than exploiting direct software vulnerabilities. The incident followed a similar GitHub repository breach at Okta just days earlier, pointing to a coordinated or trend-driven threat landscape. For CTOs, the case reinforced the critical importance of enforcing strict token lifecycle management, monitoring externally hosted repositories with the same vigilance applied to internal systems, and maintaining transparent, timely crisis communication even when the direct customer impact is limited.
Related: Why Aren’t There More Women CTOs?
How Should a CTO Manage a Crisis? [2026]
1. Foster a Culture of Continuous Improvement
a) Learning from Mistakes
Establish a structured approach to learning from past incidents. This involves promoting a non-punitive culture where feedback is actively encouraged, and failures are seen as opportunities for growth and improvement. Implement systematic tools and processes such as root cause analysis to investigate any incidents or failures thoroughlyinvestigate incidents or failures thoroughly. For instance, a detailed analysis should be conducted after a system outage to identify the exact sequence of events that led to the failure. This could involve examining system logs, interviewing relevant personnel, and using software tools to analyze the data trail left by the system before it went down. The findings from these analyses should be compiled into a report outlining what went wrong and actionable recommendations for preventing similar issues in the future. Regular review meetings must be held to discuss these reports, ensuring all team members understand the outcomes and the steps required to implement changes.
b) Innovation in Response
Encourage innovation as a fundamental response mechanism to improve systems continuously. This means staying abreast of and integrating the latest technological advancements that could prevent future crises or mitigate their impact—for example, adopting machine learning algorithms that analyze patterns in network traffic to identify and alert on anomalies that could indicate a security breach. This proactive approach allows the IT team to address vulnerabilities before they are exploited. Further, explore using artificial intelligence in predictive maintenance to forecast equipment malfunctions before they occur based on historical data patterns and usage rates. By implementing these advanced technologies, the organization enhances its crisis response capabilities and positions itself as a forward-thinking leader in technological resilience.
c) Collaborative Improvement Efforts
Foster collaboration across departments to ensure continuous improvement is a shared responsibility. Encourage departments to share their insights and challenges regularly, which can lead to innovative solutions that benefit the entire organization. This could be facilitated through cross-departmental workshops or innovation labs where team members from different business areas brainstorm solutions to shared challenges, leveraging diverse perspectives for holistic improvements.
2. Implement Robust Incident Response Plans
a) Preparation is Key
Create a comprehensive incident response plan that caters to the crises most likely to affect your organization, such as cybersecurity incidents, hardware outages, or software failures. This plan should detail each step to be taken in the event of an incident, including initial response actions like isolating affected systems, conducting forensic investigations to identify the source of the issue, and establishing communication lines for ongoing updates. For example, in the case of a data breach, the plan would specify the immediate locking down of affected accounts and databases, initiation of data tracing to understand the extent of the breach and notification protocols for regulatory compliance.
b) Team Readiness
Regularly train and prepare the IT team with drills to simulate realistic crisis scenarios. These simulations should test the team’s ability to respond quickly and effectively under pressure. For instance, randomly simulated network breaches can help assess how well the team can enact the incident response plan without prior warning. The drills should also include scenarios for less common but potentially devastating crises to ensure the team is prepared for any eventuality.
c) Tools and Technologies
Continuously equip the IT department with state-of-the-art technologies that support effective crisis management. This includes automated threat detection systems that can identify and analyze potential threats in real-time, advanced monitoring tools that provide ongoing surveillance of the IT infrastructure, and strong encryption practices to protect data integrity. These tools should be regularly updated to adapt to new threats and incorporate technological advancements, ensuring the organization stay at the forefront of cybersecurity measures.
Related: What Can CTOs Do to Improve ESG Investment?
3. Transparent and Timely Communication
a) Internal Coordination
Develop and maintain a robust communication strategy that ensures all team members are consistently informed about the crisis’s status and the steps to manage it. Utilize internal communication platforms like Slack, Microsoft Teams, or email to send real-time updates and coordinate response actions. This internal communication plan should define who needs to be contacted, at what intervals, and through what channels, depending on the evolving nature of the crisis.
b) Stakeholder Communication
Establish a formal communication protocol for engaging with external stakeholders, including customers, partners, regulators, and the public. This protocol should include predefined templates for crisis communication to ensure messages are clear, consistent, and legally compliant. During critical situations, consider providing updates at regular intervals—potentially hourly—to manage stakeholders’ expectations and maintain trust. For instance, in a service outage, provide regular progress reports on the restoration efforts and any temporary solutions to mitigate the impact on customers.
c) Post-Crisis Review
After resolving a crisis, conduct a thorough debriefing session involving all key stakeholders to analyze the effectiveness of the response. This session should identify what worked well and where improvements are needed. Document these findings in a detailed report and make it accessible to relevant parties via the company intranet or at a dedicated meeting. This documentation records the incident and the response and is a vital resource for training and refining future response plans.
4. Utilize Data-Driven Decision Making
a) Data Analysis for Insights
Utilize advanced data analytics tools to continuously monitor and assess the performance of various systems and the IT infrastructure. This includes setting up dashboards that track real-time data on network traffic, access logs, and error reports to quickly identify anomalies that could indicate a security threat or system failure. The CTO can identify recurring issues or potential weak points in the system architecture by analyzing trends over time. For example, suppose data shows repeated unauthorized access attempts from a particular region or IP range. In that case, the IT team can strengthen firewall settings or implement geo-blocking measures to mitigate those risks.
b) Predictive Analytics for Proactive Management
Leverage predictive analytics to forecast potential future crises based on historical data and machine learning algorithms. This approach involves developing models that predict the likelihood of system failures or security breaches based on patterns identified in past data. For instance, predictive models can alert to the imminent risk of hardware failure in critical servers, allowing for preemptive maintenance or replacement before the failure occurs, thus preventing downtime. Additionally, predictive analytics can simulate various risk scenarios to see how they impact the organization, enabling the CTO and the crisis management team to prepare more effectively.
Related: Evolution of CTO Role
5. Establish Clear Leadership and Responsibility
a) Define Roles and Responsibilities
The CTO needs to establish clear lines of accountability and responsibility within the IT department and other critical crisis resolution teams. This clarity is achieved by creating detailed role descriptions that outline tasks each team member is responsible for during a crisis, along with the authority levels and decision-making powers associated with each role. This information should be documented in an official crisis management handbook accessible to all staff, ideally hosted on the company’s intranet or a dedicated crisis management app. Constant training sessions should also be conducted to ensure that every teammate understands their role and how to perform it under stress.
b) Leadership Visibility
The CTO should take an active and visible leadership role during a crisis. This means being present in the crisis management war room, participating in decision-making, and communicating openly with the crisis management team and the broader organization. Their visibility can be enhanced by sending out regular updates via company-wide emails, video messages, or virtual meetings to keep everyone engaged and informed.This visible leadership helps to reassure employees and stakeholders that the crisis is being managed effectively, which is paramount for maintaining morale and trust. Furthermore, by being actively involved, the CTO can make swift decisions, provide necessary resources, and adjust strategies as the situation evolves.
6. Collaborate Across Departments and With External Experts
a) Cross-departmental Collaboration
Form a cross-functional crisis management team that includes key personnel from IT, operations, finance, human resources, and any other relevant departments. This team’s role is to ensure that all aspects of the organization’s operations are considered when planning for and managing crises. For example, operations can provide insights into supply chain vulnerabilities, finance can forecast the financial impact of various crisis scenarios, and HR can manage employee communication and support. Regular meetings should be scheduled to discuss potential vulnerabilities and update crisis strategies accordingly. Crisis management software tools can also enhance communication and coordination among these diverse teams.
b) Engagement with External Experts
Establish strong relationships with external cybersecurity firms, industry experts, and academic institutions. These partnerships can provide crucial knowledge and insights during a crisis. For instance, cybersecurity firms can offer advanced threat detection services and immediate support in the event of a breach. At the same time, academic institutions might provide the latest research on risk management or recovery strategies. Engaging regularly with these experts through workshops, joint exercises, and advisory sessions can keep the organization updated on the current developments and best practices in crisis management.
7. Develop and Maintain a Business Continuity Plan
a) Comprehensive Coverage
Develop detailed continuity protocols for each critical business function to ensure minimal disruption during crises. This includes establishing alternative processes like remote access for key personnel if physical office spaces are unavailable, redundancies in critical IT systems to prevent data loss, and flexible work arrangements to maintain productivity during external disruptions. Each department should have tailored continuity steps that align with their specific operational needs.
b) Regular Updates and Testing
Continuously update and test the business continuity plan to reflect changes in the business environment, technological advancements, and new potential threats. These tests should be conducted at least bi-annually and involve realistic scenarios designed to evaluate the readiness of different departments and the effectiveness of the communication and coordination strategies in place. The results from these tests should be reviewed thoroughly to identify weaknesses in the plan and provide targeted training or resources to address these gaps.
8. Prioritize Employee Training and Awareness
a) Ongoing Training Programs
Implement a continuous education program that covers cybersecurity, emergency response, and crisis management best practices. These training programs should be mandatory for all employees and conducted regularly to ensure everyone knows the latest threats and how to respond effectively. Training can be delivered through online courses that enable employees to learn at their own pace and in-person workshops that offer interactive learning experiences, such as role-playing exercises or group discussions.
b) Crisis Simulation Exercises
Organize biannual simulation exercises that mimic real-world crisis scenarios. These exercises should be realistic and require employees to react in real-time. This hands-on approach helps to test the organization’s operational response capabilities and allows employees to practice their roles under pressure. After each simulation, conduct a detailed review session to discuss what went well, what didn’t, and where improvements can be made. This feedback loop is critical for refining the crisis response strategy and enhancing overall preparedness.
9. Strengthen Security and Risk Management
a) Regular Security Audits
Implement a schedule for semi-annual security audits through reputable third-party firms. These audits should thoroughly examine all aspects of the IT infrastructure, including network security, application security, and data protection practices. The objective is to identify vulnerabilities that might not be apparent from internal reviews. For example, third-party auditors might employ advanced penetration testing techniques to simulate external and internal attacks, testing the effectiveness of both physical and cyber defenses. The audit findings should be documented meticulously, with clear recommendations for addressing any identified issues.
b) Risk Management Framework
Develop a dynamic risk management framework integrated into the organization’s daily operations. This framework should include systematic risk assessments performed regularly or when significant changes occur in the business or technological landscape. It should detail the processes for identifying, assessing, and prioritizing risks on the basis of their potential impact on the organization. The framework should also outline mitigation strategies, including preventive measures and response plans. To keep the framework effective, it should be reviewed and updated regularly to incorporate new insights, emerging threats, and technological advancements. This living document ensures the organization remains agile and responsive to changing security landscapes.
10. Leverage Technology for Real-time Monitoring and Alerts
a) Implementation of Monitoring Tools
Deploy state-of-the-art monitoring software across the IT network to maintain a continuous check on the health and security of all systems. This software should provide a comprehensive dashboard that IT staff and decision-makers can use to view real-time data about network traffic, system performance, and security alerts. For instance, the dashboard could display key performance indicators such as server load, network latency, and the number of active connections, alongside security alerts like attempted breaches or unusual data transmissions. By monitoring these metrics, IT teams can quickly detect and respond to anomalies before they escalate into more serious issues.
b) Alert Systems
Establish a multi-tiered alert system that categorizes issues by severity and escalates them accordingly. For example, low severity alerts might warrant an email notification to the relevant technician. At the same time, high-severity issues might trigger an immediate SMS or phone call to top-level IT managers and the CTO. This system ensures that critical alerts are promptly addressed by decision-makers with the authority to mobilize resources and coordinate a response. Additionally, integrating this system with mobile technologies allows for real-time notifications to the relevant personnel, regardless of location, ensuring swift action can be taken even during off-hours or when key staff are away from the office.
Conclusion
The role of a CTO in crisis management is pivotal in safeguarding an organization’s technological and operational integrity. A CTO can effectively mitigate the impact of crises by implementing robust incident response plans, fostering transparent communication, and promoting a culture of continuous improvement. Moreover, the strategic integration of advanced technologies and cross-departmental collaboration further enhances the organization’s resilience. As organizations navigate the uncertainties of the digital age, the CTO’s ability to lead through crises protects the organization and sets a foundation for future growth and innovation. In essence, adept crisis management by a CTO is not just about replying to immediate threats but also about building a resilient framework that supports sustainable success in an ever-changing technological landscape.