What's Taking IT So Long To Fix the Problem?

Posted
By: Eric Wolferman We recently tangled with a nasty problem in our classified-advertising-system network that took us more than two weeks to wrestle to the ground. Groups of PC workstations were randomly losing their connections to the server. We were able to reconnect them without too much trouble, but the problem continued to plague us. Needless to say, it was very annoying for users taking ads.

As we painstakingly worked to identify and isolate the problem, users and their managers pleaded for relief and questioned why it was taking so long to solve the mystery. Ultimately, we determined we were fighting more than just one problem, but some changes to a central switch in the network finally seemed to put us right.

Nevertheless, the question posed by system users is legitimate: Why was it so difficult to track down the culprit?

The answer lies in the growing complexity of computer networking -- not only at newspapers but in all businesses. A better understanding of the parts under the hood can help frustrated executives understand what is happening.

In the 1970s, computer systems consisted of "dumb" terminals connected to a central brain, the mainframe. Communication and interaction between the mainframe and the terminals was straightforward -- cables generally connected the workstations directly to the main box. When things went wrong, there were a limited number of places to look for the problem.

With the introduction of personal computers in the 1980s, a new scheme was needed to enable these self-standing machines to talk to each other. This new scheme would have to allow communication among many different brands of machines running a variety of operating systems. The solution was the emergence of a number of standards, or protocols, that spelled out how computers would send data to each other.

Several companies, such as Banyan Systems Inc., IBM, and Novell Inc., introduced solutions for networking PCs and pretty soon such networks were commonplace. They all used basically the same approach: a special card in the PC to allow communication with other machines, hubs and switches to provide central points to plug things together, and lots of wire to connect it all.

Since then, this simple scheme has been refined and improved to the point of allowing thousands of workstations to communicate on a network. In fact, the celebrated Internet is simply a grand extension of this technology, tying together computers from around the world.

But as network technology became more sophisticated, the number of pieces and parts multiplied like rabbits -- and the components became vastly more complex. Whole departments are now dedicated to the care and feeding of networks that are critical to our ability to do business.

A hundred miles of bad road?

Our network staff estimates we have installed upward of 30 miles of standard networking cable in one of our facilities since 1996. In addition, we maintain nearly 50 separate computer servers and another 60 miles of telecommunications cabling in the same building. Multiply that by two operational centers, two printing plants, and assorted district offices -- and you begin to realize the scope of what we are dealing with.

Hundreds of workstations, connected to various servers through a growing network of switches and routers, direct communications traffic throughout the enterprise. To further complicate matters, the networks in our facilities are tied together via data-transmission lines. It is little wonder that diagnosing problems and isolating errant components can be as difficult as mapping the human genome.

Networks also tend to grow wild from time to time. As workstations and servers are added, we often meet immediate needs by adding another component or changing some configuration -- with every intention of returning in the future with a more permanent solution. Meanwhile, the Band-Aids mount, making it even more difficult to trace problems when things go awry.

If all this sounds a bit defensive, well, perhaps it is. But the technology industries must recognize the growing complexity of network science and present solutions to simplify the chaos. There are lots of new tools to monitor and trace network problems, but we are still far from foolproof methods to quickly identify network failure points and correct them.

Computer systems have evolved like Tinkertoys. Having pieces and parts that fit together to suit many different business environments provides great flexibility and, in many cases, significant economies. But the price we pay, at least for now, is the ongoing burden of keeping the swelling miles of data highways reliable and stable.

Comments

No comments on this item Please log in to comment by clicking here