Showing posts with label IT Governance. Show all posts
Showing posts with label IT Governance. Show all posts

09 October 2020

Not everyone should be an Internal Auditor

Sometimes Internal Auditors shouldn’t be Internal Auditors. Sometimes the role can be, no matter how much effort is expended to avoid this, confrontational or with the potential for conflict with the auditee (and others). This is particularly the case when there are strong personalities on the ‘other side’ of the audit process. I ran into exactly such a situation, as I’m sure have most of us. Remember, however, that just because someone is not appropriate for Internal Audit that does not mean that they may not have a lot to contribute to the business.

A number of years ago, I was engaged by a bank to perform a number of IT Audits. The bank had a full Internal Audit function but only three IT Auditors. The audit programme, however, included too many audits to be completed by the team that was available (for various reasons, only one of which was to too much work for the available resources).

After cutting my teeth on a couple of simple reviews, the Audit Director asked me to take a look at the implementation and use of the Project Management Methodology in a couple of the major projects that were in-flight at the time. These were significant projects, being run by and for different parts of the bank. Each had external project managers, and each seemed to be running to time, budget and promised deliverables. There were no particular reasons to worry about the projects.

Enter Bob (not his real name), a somewhat meek Internal Auditor, who chanced into IT Audit from a role as a bank branch auditor. I had worked with Bob before at another institution, and knew some of his strengths and weaknesses.  The Internal Audit Director said to me “I’d like Bob to work with you on this audit”. Really? Well, okay. “It will be good for him. He’ll learn something, and hopefully will become a better auditor.” He saw the horror in my face.

“I really need you to do this, but let me know how it goes”.

So the audit began. Each project provided all the requested information, and both were open allowing interviews with key project personnel and the projection managers. The project sponsors were comfortable the progress, and the user communities were looking forward to the new systems and processes, even though these were months away.

The projects were running smoothly, and the audit did not find any unreasonable budget to actual variations, or undue and unexpected slippages in estimated deliver dates, resource requirements, etc. Risks were documented (inadequately, but there was some consideration of risks). Of course, the primary purpose was to confirm the implementation and use of the corporate-mandated project management methodology.

While everything is going smoothly, a finding that process is not being followed can be a difficult finding to make and defend, especially when the processes will add effort and probably increase the resources and costs required to accomplish the project or set of tasks.

Add to that the personality trait of many good project managers – a straightforward manner and an air of confidence that can be used to ‘encourage’ focus on goals. They are confident, and they exude confidence, and that is one of the ways that they provide comfort to stakeholders, encourage teams, and deflect or reduce potential conflict or disagreement. This sometimes can manifest itself as arrogance and bullying.

And we faced two of these individuals. They had the backing of their respective General Managers, they were confident, they were delivering, and they really didn’t need Internal Audit second-guessing how they were going about achieving their missions.

I sent Bob to carry out some interviews, collect documentation, read it and summarise his thoughts. We talked through what he was seeing. We combined our work and work papers, and we arrived at our conclusions. We wrote up the draft report, and prepared for the exit-interviews with the two Project Managers. 

As the fieldwork progressed, Bob became more and more agitated, and at times seemed distracted. Finally, with the fieldwork completed and the draft report ready, we scheduled the exit interviews. Twice.

Then a third time, with each of the other two being cancelled and rescheduled.

Finally, the day arrived. I arrived in Internal Audit, and seeing Bob, said “Fantastic, today is the day. They’ve not cancelled or postponed. We’re ready.”

I looked closely at Bob. “Are you alright? You look tired.”

“I haven’t slept all week, I’ve been so worried about this meeting” was his response. Worried? Why? All our ducks were in a row, all the documentation was completed, the draft report was written, the findings reviewed, and the key points ready. All that was needed now was a conversation with the PMs, and to give them an opportunity to take the draft back with them and write up their comments, responses and action plans.

Focusing on the coming meeting, I put his comment away in the back of my mind, something for later.

We had our exit meeting. We outlined the audit, the fieldwork performed and the data and information reviewed. We presented our findings. The PMs read the Executive Summary, looked at each other, and after a few questions said “You’re right, we use our own methodologies. They are not the corporate-approved methodology. We will talk to our teams about how we will implement and use the standard methodology. We will need to train our people, and we might need some training also.”

Done. 

Yes. It was that ‘easy’. The data was there, the documentation was there, and we did not attack their methodologies or pick holes in what they were doing. We were not auditing the effectiveness of their personal leadership, and we were not questioning the performance of the projects (although we did look at status reporting, steering committee reporting, budgets to actuals, etc). We had a specific scope and we audited to that scope, cognisant that other issues may come up.

What I didn’t expect was that the primary finding of serious concern was that one of the auditors was not able to perform the audit. Having worked with Bob in the past, it all came together then. He simply was not capable of assertive support of any position. His default in any potential conflict was not to address the issue, but to seek someone who could deal with it on his behalf.

When all was done and the report was issued, I stopped by the Audit Directors office. I told him what had happened, and said I was deeply worried about Bob, his mental state and his fitness to be and Internal Auditor. Furthermore, there was the very real potential that Bob would bring Internal Audit into ‘disrepute’ within the bank by not being adequately assertive or able, when pushed, to deal with highly assertive individuals. In the worst case, such an auditor might miss a critical control and technical issue, or fail to push for acceptance and resolution of a critical weakness, potentially endangering the bank itself. The IA Director knew we had worked together in the past, in fact, all three of us has been at another bank at the same time in the past. He “inherited” Bob when we took over IA in this bank. He knew what he had, but there was little he could do directly.

We talked, and eventually, I said “You have to get him out of Internal Audit. He will have a nervous breakdown, or worse. This is not the right job for him.” The IA Director agreed and asked for my suggestion. My view was that Bob had a solid knowledge of retail banking, adequate IT knowledge, and understood both the bank and the banking sector. Firing him would only compound Bob’s issues and would be wasting an otherwise perfectly decent person and skill-set. “Find him another job in the bank. For you and for him”.

Checking in with the IA Director a couple of years later, I asked what was the final outcome with Bob. The news was all good. Bob was encouraged to apply for, and was appointed to, a role in the Retail Product Development team, and was to all reports thriving. Conflict was not an issue, because he was supporting product developers who were, by nature, positive and had the support of the executives. His knowledge of the bank and banking products served him well.

Most of all, a ‘wrong fit’ was rectified, and IA was seen as a potential source of good quality people for the business, and not tarnished as the home of people who were not able to provide the challenge actually needed in healthy organisations.

What are the attributes of a good Internal Auditor? There is a long list. Near the top of any list must be confidence in the correctness of the principles that the auditor is espousing; of effective control, process effectiveness, risk identification and assessment, and confirmation by the auditee of the findings and potential impact. Meekness is not a desirable attribute.

  

05 March 2019

IT Audit - sometimes you need to escalate

A common facet of contracts is a true-up clause that pushes a disagreement on price or capacity into the future, with actual usage or consumption to be calculated at a future date or time. Think of the classic French Bistro (in the outback of France, no in a London or New York suburb), and the bottle of house red that is automatically delivered to your table. Or the bottle of whiskey in the officers mess in the Indian Raj, with the line drawn on the bottle. When the meal is finished, or the drinking is done, a new line is drawn, and you are charged for the difference - the amount consumed.

There is no contract that requires you to consume the entire bottle(s), or a guarantee that you will only drink three-quarters. The contract is settled at a later time. The core of this contract is that all can clearly see what was consumed, and there can be little dispute as the actual quantities and therefore the final bill.

I have seen computer systems contracts with just that type of resolution built into the contract. 

Many years ago, I was asked to look at a contract that had such a true-up clause in it. The computer vendor had estimated that a certain level of computing power (mainframes) would be required, while the client estimated a lower amount would be required. In the days before on-demand cloud infrastructure, computing power came in "boxes" of defined "MIPS"(Millions of Instructions per Second - a quaint concept to us today). You got the whole box, or no box. The vendor believed that a certain number of "boxes" would be needed, while the client thought otherwise.

The system was of too much importance however, to allow for the implementation of inadequate computing power, and so both partied agreed to install enough to ensure smooth functioning. The vendor was adamant that their estimates were right, so insisted that the total amount of processing power be installed.

Through the negotiations, a final difference of $18 million was arrived at, out of a total contract value of approximately $80 million. The parties agreed then, as is not uncommon, to split the difference three ways.


  1. The client agreed to pay $6 million.
  2. The vendor agreed to discount $6 million.
  3. The parties agreed to review system usage at the end a year, and split the remaining $6 million based on the actual usage.


Makes perfect sense, if the actual usage can be measured and recorded, and if monitoring and system optimisation are in place on the client side. Like the line on the bottle, the utilisation level could be measured and a line drawn across the capacity of the systems.

Unfortunately, the client failed to put in place the monitoring. As a former mainframe systems capacity planner, I knew what monitoring would be required, and I knew exactly how the vendor would demonstrate that the application actually did require the full amount of computing capacity that was originally estimated. I had, in fact, worked for that vendor.

As the IT Auditor, I recommended that the monitoring should be put in place, and provided guidance on what and how to perform that monitoring. I also recommended that such monitoring should be performed on an ongoing basis, so that management could track how much of the $6 million they would "owe" at any given month-end, so that system optimisation could be performed. 

Nothing happened.

Again, in three months, I recommended that the monitoring be put in place. Again nothing was done. All the while the clock was ticking down to the performance date, and it was looking like the $6 million would be owed to the vendor.

Having received no response from the CIO, and in fact, having been told by the CIO that Internal Audit really didn't know what it was talking about, that Internal Audit knew nothing about IT, and that IT auditors were a particularly incompetent bunch, we felt there was no option but to escalate. A one-page memo was prepared and sent to the CEO (the same CEO who sent a two-page memo to all managers telling them that all correspondence to him should be in one-page memo form) outlining quickly the situation, and the (lack of) response from the CIO.

The result: After an independent review of IS's work lasting all of one day, the CIO was fired, and new negotiations were opened with the vendor, and a pre-emptive agreement was reached that saw the client pay the vendor $3 million. The vendor forgave the other $3 million.

Ultimately all agreed that they would not be able to draw a line on the bottle that each party would agree to, so it would have been almost impossible to agree exactly how much had been consumed.

But failure to implement basic monitoring and management cost the company $3 million that they should not have needed to pay. 

19 October 2018

Watch the Backup Tapes - Yes, Really

A few weeks ago I was standing in an office when one of the IT people walked through. He was on his regular walk to deliver backup tapes to the "off-site" location, and this also gave him the opportunity see if he could trouble-shoot an issue for a senior exec. And there they were, backup tapes, just sitting unattended on a desk, as Mr IT was at the other end of the floor looking at the laptop. Yes, the data could have been sent across electronically, and certainly, the company's major systems are all automatically backed-up with almost constant checkpoints. Those systems can be flipped from the primary production to the back server farm in a matter of minutes.

But back to the backups; what was on those tapes? What would have been the impact on the company if the tapes disappeared? Source code? Customer data? Test scripts? Downloaded movies? Confidential emails?

After all, the tapes are a standard format, and the operating system is standard, so there would have been a minimal challenge to recover everything from those tapes. 

Sometimes we see "non-critical" systems continue to be managed as stand-alone environments, disconnected from the corporate environment, and most especially from the automated backup world.

I was tempted to pocket one of the tapes and see if he noticed it was missing. As there were only two tapes, or the cartridges that replaced "tape" decades ago, I was pretty confident that he'd notice. Cheap joke, not worth the effort.

But it did remind me of a blatant demonstration that I had to give a General Manager almost twenty years ago. Because sometimes a such a demonstration is what is needed to make the point. This is especially true with backup tapes.

We were (internal) auditing a subsidiary of a telephony company, and were in their Auckland office. The subsidiary was building a new product for corporate clients, enabling the clients to take a single file of all telephone data, already data-cubed and with a set of associated software to allow multi-dimensional analysis of that data. All very cool for the time. 

The systems ran on a small set of servers, actually just very powerful under-desk PCs, up-configured about as much as was possible in that day. And as this was a subsidiary and not integrated into the telephony company, the servers were of course in a room in their offices in the Auckland office building. And of course, one of their main programmers also worked in that same cubbyhole of an office.

Furthermore, the backup tapes were kept in the drawer of the desk next to the computer. The door to that room was of course locked, I was told, even though the door stayed open all day. At night, when the programmer left, he closed the door behind him, and then unlocked it to get back in the next morning.

By the end of the audit, we had our draft report and findings completed, and we were ready to present these to the General Manager. The only time slot he had was after hours, so we sat down with him at 7pm to go through the draft report and findings.

If there is one thing that I've learned about auditing, it is that the enthusiastic nodding of the executive is as frequently faux-agreement so that you will just go away as it is agreement to fix the issue, no matter how trivial. In fact, the more trivial, the greater the chance that the nodding will actually be a "please go away" message.

We were not getting any of those messages until we got to the IT Audit findings, and then the nodding started. The issues were outside his area, and he didn't understand the issues, their severity or the what effort would actually be required (all issues had been cleared with the IT people first of course). So he entered "thank you for these very helpful findings, please go away" mode.

Being fair, most of the issues were fairly minor, but I knew that a demonstration was needed. Something to make sure that the next morning, he would actually call in the IT people and get the issues, if not resolved, at least on the agenda for resolution.

When it came time to talk about backup security, I stopped, stood up, and excused myself from the room. "Pardon me, I'll be right back", and walked out...

...down the hall to the IT office/cubbyhole...

...Opened the door (that I had unlocked earlier in the day, guessing that the programmer would simply pull the door closed behind him when he left), opened the desk drawer, and walked back down the hall to the General Manager's office...

...where I put a box of tapes/cartridges in the middle of the table and said "these are your backups. If these were to fall into the hands of our competition, this product will no longer be commercially interesting, as our competition will be providing the same service within a month. Furthermore, if there is a fire or other reason this floor is not accessible or damaged, these tapes and your primary server will be lost."

I had to explain to him that the tapes were in a box where almost everyone could get them, and it would have been frighteningly easy for someone to simply replace a tape with an empty one, and walk out with a complete backup of the code, customer data, and effectively the entire product.

They were also kept right next to the machine that they backed-up. Lose the room, lose the computer, and lose the backups at the same time.

Of course, the IT guys weren't thrilled with me, but they suddenly did have the budget to install an electronic back-up to the main telephony company's computer centre and backup environment.

The lead auditor on the job was also shocked, though mostly because I hadn't warned him.

It is easy to become complacent, thinking that technology has come so far, that the silly things we allowed to happen years ago cannot happen today. Yet there is another less we should take; IT governance is first and foremost a people driven set of processes, not technology processes - technology allows us to make those processes work more efficiently, but they remain human processes. And no matter how much technology has progressed in the past twenty years, the human remains fundamentally the same.


13 September 2018

Focusing on SPOFs is never a WOFTAM

If there is one thing I really enjoy, it is having fun with acronyms. I did make one and tried to get people to use it but to no avail. I think it was an example of self-WOFTAM (defined later). But as this is not a blog about fun with language, but more about Risk Management and other Random Comments, I'd like to talk about SPOFs.

SPOFs are closely related to that other topic I have covered, the Delegations of Risk Authority, if only because so frequently when someone says "we've accepted that risk", you can be there is a SPOF in there that they simply unable to justify correcting, that they do not know how to correct, or that they have had too much push-back to keep trying to solve.

Too often there are invisible Single Points of Failure (SPOFs) across an environment. These naturally occur due to the constant changing of configurations, slowly introduced incompatibilities, and sometimes (too often) by short-cuts built into environments to keep the cost of the application to within the level of spending that was thought to be acceptable.

The following tale is, in principle and flow, a true tale.

I sat with the Systems Architects and looked at the environment diagram they gave me, and asked them to explain how systems and applications were reached by our users, and how a critical outage could have lasted so long (almost a week). The conversation when something like:

"Well, the application is on this server, and since the users are in this building, we keep that server in the network room in the same building."
Good so far.
Okay, where is the alternate server or backup system?
"We have a backup server in the network room in the other building."
Even better.
What happens if this server room should go down - that's never happened, has it.
Sideways looks from person to person.
"Well, it did go down last year."
And?
"We brought it back up."
Okay. How long was the system down?
"About a week. No, four days. We're pretty proud of that, it was only four days, and the users were able to keep working off-line during that time."
Weren't these users customer facing?

"Only some of them and they were able to take details on paper and ring the customers back. After the system came back."

SPOFs

It turned out that there was not one, but a number of Single Points of Failure (SPOFs) across the architecture, each waiting for that wonderful moment to manifest.

First, the primary server was in one location, and the backup server was in another. Backups were taken every night, and then transported to the backup site. Tapes were cycled between the sites. All pretty standard (and now replaced by a link for online backup to a remote backup device), except that recovery on the backup server to the copy of the application had not been tested.

So why wasn't the application just brought back up on the alternate server?

SPOF number two. Well, now it gets interesting. The primary server had a physical fault that required the replacement of physical kit, and the backup application was loaded on a server which acted as the primary server for an application with a different operating system release level and patch history, and therefore did not have a compatible operating system configuration. The "backup server" had not been updated, and was missing key licenses for underlying software required by the application.

The first attempt to bring the backup server into production failed because the underlying operating system and supporting software were not up to date. That took a couple of days to correct, with a rebuild of the backup server, while ensuring compatibility with the primary application on that server. Meanwhile, the primary service was being worked on by the engineers, who, on the recovery of the server, discovered that the network connections to the alternative site would not allow a pass-through for the users.

The list of SPOFs in this situation continued to grow, each needing to be worked through or around.

SPOF number three. The lack of adequate bandwidth required building a new set of IP pipes that would enable authorised users to access the back from their primary location, without opening the application to access from any IP address. While not difficult, it was time-consuming.

Further adding to the problem was that as each SPOF was worked around, the primary objective of recovering the application meant that the SPOFs were mentally abandoned. Maybe one day we will go back and look at them again, but for now, our only priority is system recovery.

This was a wake-up call. A string of SPOFs had almost crippled a key element of the business.

What next?

The identification of SPOFs is not easy, and is a project in itself. In addition and from experience, once all SPOFs have been identified, the probability is that only 80% have actually been found. For months after, expect someone to come to the project team or lead and say "Um, I think we might have found another".

The list of SPOFs then needs to be reviewed, ideally by a combination of the technical people, architects and with input from users (for confirmation of criticality), for:

  • Completeness,
  • Criticality,
  • Likelihood,
  • Cost of remediation, and
  • Interdependence (caused by or contributing to another SPOF).

From this, a plan for remediation can be developed. Bearing in mind that such plans should start with remediation of the most critical and the most interdependent SPOFs. Eventually, the cost against remediation benefit break-point will be reached. However, it is not the role of the technical team to determine that cutoff. The cost of remediation needs to be determined, and a multi-phase and costed project plan developed. Multiple scenarios of levels of remediation at various price points should be provided so that those who have the authority to approve spend and the authority to accept a level of residual risk have the information that they need for decision-making.

Too often I've seen the "we've accepted that risk" type of response when considering specific SPOFs or elements of the plan. It is not the role of the technical team to accept those risks, but to communicate the risk and the cost of remediation to those with authority to accept that residual risk.

A note on cost.

It is almost impossible to remove all SPOFs. First, the costs become higher than the potential cost of a resulting event caused by a SPOF. Second, there are many SPOFs that, through analysis, will be seen to be of such low possibility that the cost of remediation will probably far outweigh the cost of a mad-scramble to resolve the situation should that SPOF eventuate.

Special consideration should be given to the high cost (for marginal return) and long time frame for remediating SPOFs. For the higher cost elements, it might be reasonable to identify human and process workarounds in the event of the failure of that specific SPOF. For longer duration remediation elements, additional consideration should be given beyond a simple cost/benefit.

For example, high capacity network links can take some weeks to be installed, and the lack of such links can turn a simple recover effort into a Business Continuity and Disaster Recovery exercise.

Finally, the Future.

Even when the SPOF remediation plan has been approved, work has taken place, new equipment installed and tested, and the project has reported back to the steering committee that the project has accomplished its objects, within budget and within time (no comment), the management of SPOFs is not done.

As mentioned above, systems and environments evolve. It will not take long for divergences in system configurations to creep in, for levels of installed software to become out of synch between production and backup, or for model office environments to slip out of synch with the production environments that they are meant to replicate.

A SPOF review on an annual basis will not be a WOFTAM, and will identify new SPOFs, or may result in a reassessment of the importance of a SPOF that was previously accepted.


(WOFTAM: "Waste of F*** Time And Money" - do feel free to use that, as it is one acronym that I've found myself muttering under my breath for years. Oh, and there is no copyright on it.)