Tuesday, February 03, 2009

Out of the bubble...

Years ago I had a friend who as a psychologist who used to work among media people in Southern California. One of the phenomena he talked about which rang bells with me was the 'bubble effect' among media people. When they were working on a project they went into a bubble and were not like normal human beings, emerging as [relatively] normal people from the bubble when the project was over.

As a result of the attacks we had in October 2008 we decided to upgrade two of the servers. These servers would be three years old this January. They have a life span of approximately three years before needing an upgrade. Each time we upgrade, it's not just a case of new hardware and put in the CDROM and type install but we have to think through of the needs, particularly security needs, for the entire life of the server. So, Peter and I thought through what we felt we needed for the next three years.

The last time we did an upgrade it took Peter and I approx one month - ie two man months of labour. This time we then brought a colleague over from Egypt and between us we hoped that in one month we could get the new hardware and software working to replace the old. We expected that three man months would be about right... but it wasn't... it has taken approx nine man months of labour to do the upgrade!

Why so much more? We were adding many extra security features which proved very much more complex than we expected. In fact, the added security was somewhat frightening for me. I was thinking back over the steps to get here... from the original servers... to the upgraded servers... to these new servers. The complexity seems exponential. The upgraded servers were about twice as complex as the original, and the new ones about four times as complex as the upgraded ones. Hence I'm already getting edgy about what it will be like in time years time... sixteen times as complex as these new servers?

The main security issue is for each application within the server to be isolated from every other application, running in a virtual server with its own security. That way an attack on one part should not affect the whole. So in reality it's like going from two servers with eight primary applications to having eight servers with one major application each. But they cannot be totally isolated and we then had to work out secure ways that each application could talk to the others that they needed to. Yipes... yes, horribly complex and hence why I was concerned about the future.

Over Christmas we had another colleague and his wife over so that he could have extra training and to plan together the next step of the project he is working on. So... having been in the upgrade bubble and not completely out of it, it was straight into another bubble. Not that it was bad, but it did mean we didn't get a break at all.

We also had end of year calculations to do and create budgets and plans for 2009. Actually doing this took Pete and me out of the upgrade bubble for a while and did enable us to see the 'wood for the trees' which was helpful. But budget planning is not one of my favourite pastimes. Just before the end of the year Peter remarked, 'You know, I wouldn't do this if they paid me...' I had put into works exactly my thoughts!

Then we discovered that for various reasons we had to upgrade one of the other servers that is only one year old. We lease the hardware, so strangely enough because of the drop in price of hardware the new server will be under 2/3 the price of the old one! We had to wait for delivery of the hardware which was handed over to us yesterday. Yes, that means we are still in the upgrade bubble.

One Egyptian colleague is still working with us on the upgrade process, which we hope to complete by the end of February. It should be quicker now on the extra server as we know roughly what we want and can copy the two we already have working.

So does that mean its all straightforward for a while? No, not entirely... tomorrow Peter is off to Holland for five days to attend a conference and then next Wednesday I fly off to Australia for about a month, partly to attend a conference, partly to evaluate some new software and to see if a partnership with an Australian group will happen and partly to visit other organisations and... but... its not coming together easily...

A couple of days after I had finally booked my tickets I heard that there will be a delay on the new software, which means I shall probably have to go back to Australia sometime later in the year. It's both a huge expense and a huge cost in my time. I am not best pleased to put it mildly. There is only one light at the end of the tunnel as far as the trip is concerned. If everything works out I shall see my son for a couple of days on the way back through Manila.

So there we are... it almost seems like we cannot get out of the bubble, but the bubble is expanding to keep us inside it!

Thursday, October 30, 2008

Most men, so I'm told, are glued to the TV when either the football or Olympics is being shown. Not me. I'm happily fairly oblivious to either. But now... it's the Volvo Ocean... and why do I mention it? Well, my team is currently in the lead. And, to make things even better, they have just beaten a world record.
Torben Grael and the crew of Ericsson 4 swept into the history books yesterday as the first monohull to breach the 600-mile barrier in 24 hours. They’ve been chased by men, machines and the elements in the last 48 hours – and nothing has touched them.
They had been lying fourth behind but battling it with the leaders - Green Dragon, Puma and Telephonica Black. But its pretty dreadful weather they are sailing through as Mark Chisnell puts it:
In their foaming, boiling, 25-knot wake the fleet lies scattered as the devil and the deep blue sea picked off the hindmost one by one – the cold front sweeping over them with a mix of murderous squalls and ugly waves in a pitch black night. We’re almost down to the last man standing.
If you're as gripped as I am you can follow the race online, even through a 3D virtual simulator, where the boat's instruments, when they are working, relay everything via satellite to your computer at home... almost in real time. But they don't always work. In fact, Ericsson 4 have equipment failure now.

So back to reality for me... over the past couple of weeks we have been battling murderous squalls on the technical front. Three weeks ago I wrote about the DDoS attack. One of the outcomes of reviewing this was a decision to upgrade two or more of the servers. They are three years old now and so replacing them is about due. But its not just a case of copy the files and off you go... it will take about three of us at least a month to move everything over and upgrade all the systems on the new servers. A very big job, which is why we only try to do it every three years!

Having decided to do this we brought Raed over from Egypt to help and then ordered the new hardware. We lease the servers rather than buy them, leaving the leasing company responsible for the hardware maintenance. On Monday they will pass them over to us, with a bare operating system on them and we will start the task of checking them and installing all the systems and moving the sites across.

In between all this the attacks have continued - like a cold front sweeping over us. We watch the attackers in real time, and have defense mechanisms set up to rebuff them. But trying to second guess their moves is difficult, so we have set up what is called a 'honey trap' to try and lure them in to showing their methods. This will give us some indication of how much they know about us and why certain sites are more attacked than others.

One of our partners - with a site for central Asia - was online chatting with me today and they want to increase the facilities, to start online broadcasting to their region. Another site - for the Middle East - will have new facilities and a new design before the new year. A further new site - also for the Middle East - should be live before the new year. So it feels like a 'foaming, boiling, 25-knot' race downwind barely in control of what is happening. I am looking forward to Christmas - which I hope will be the end of this leg of our race and the sites and new servers will all be behind me.

Friday, October 03, 2008

DOS attack

Most of day yesterday we suffered what was called a 'Distributed Denial of Service' or DDoS attack. This meant that web sites on one server were unavailable at times. The problem will have shown itself as either the server appearing to run slowly, or unavailable or problems within the website that looked like a MySQL problem.

So what is a DDoS attack? Well in our case all of these were caused by a whole load of computers sending invalid file requests many times per second - or at their slowest many many times per minute. What this did was to start extra instances of the web server to respond to these requests, till the server ran out of resources and failed to deliver. Normally the 'load of computers' are Windows computers with viruses [usually called a botnet] that allow them to be controlled from a master computer or robot system. All automated. Against us.

Peter eventually wrote a new rule into our automated response system to stop this happening by blocking users who try the same method of attack. Within seconds they were being blocked.

Fortunately it was a relatively minor attack. We recorded only 59 computers attacking us from the time we turned on the rule in the automated response system to block them. Today this has dropped to a trickle of 26 still attacking us in the first 8 hours of the day - all being blocked. Some botnets are huge - for instance, this August the Dutch police shut down a botnet of approximately 100,000 [Windows] computers infected and controlled by two people.

Oh, the the problem on Wednesday turned out to be a faulty cable. How come a faulty cable did all that? Well, the switch connecting to a workstation in the office, which, by the way, was turned off, sensed something strange on the cable and decided to keep trying to sort it out many thousands or millions of times per second. It also decided to tell the entire LAN about the problem [a broadcast message] again many thousands or millions of times per second. This broadcast message affected other switches and affected the server. Cable fixed, fault disappeared!

In case you're thinking that sounds rather like the DoS attack we suffered, it was. It was a type of DoS attack. The difference being that one is accidentaly, but from the evidence in the logs we can see the other was malicious.

Wednesday, October 01, 2008

Yikes its a bridging storm?

Today is a public holiday... so I should be off. I had hoped to go sailing with a friend.

Alert on my phone: All the connections to the FUP system are down... in fact sarah is down [sarah is the name of one of our servers]. This needs urgent attention. So I speed to the office.

It appears that one of the transceivers on one of the routers has failed. So I change it. No difference... but one part of our system starts to sort of work. So I check all the cables... and find that some that I need to know what they are are not labeled. [We have copious free time for labeling... not!] So I label all the cables, plug in the critical ones and everything looks fine.

I plug in the rest and... one of the servers has totally locked up. What? Crazy... cannot happen. Spend next few hours sorting out the server and everything looks fine... for a while... but my notebook cannot get an IP address. Why? So I unplug all the non-critical cables and... my notebook gets an IP address. Everything looks fine.

I plug in the rest and... one of the servers has totally locked up. What? Crazy... cannot happen.
OK, this time I learnt my lesson. I leave all the uncritical cables out, reboot and sort out the server and leave for home [dinner time now].

After dinner... I get an alert. One of the servers is not connecting. So I go back to the office... and find that in all my plugging an unplugging one of the cables has become lose. So fix it and plug in and go home.

Then the strange bit. I speak to Peter. He esplains [hope I get the jargon right] that we may have a 'bridging storm' going on. Basically its this... we have more than 50 devices [servers, routers, phones, workstations etc] in the office on 3 different physical LANs [ie networks] connected to about 16 'switches'... connected to 2 Internet connections to the outside world.

Switches are the things that connect all the devices together and talk to each other making a tree with one being the 'boss' [I'm sure Peter had a more technical word for that]. And the master switch talks to all the others telling them where in the tree they are and how to behave. If one of them wants to become the boss then an argument starts and can result in a bridging storm where some switches [and thus devices] are cut off. Why so many switches? Well, three reasons - firstly it's difficult and expensive to cable every point from a central location, secondly we have many extra points we need for testing and research and development and finally manufacturers [including manufacturers of VOIP phones now add a switch in the back of their devices.

So how would this make servers lock up? Well... we have some clever software in that to make sure that either the main or the backup is up and working. This is roughly the same language as the switches talk and maybe, just maybe, the bridging storm makes this go really crazy. Well... it makes me go crazy anyway.

Monday, September 29, 2008

A duck paddling upstream?

Another month has gone by and I am re-reading what I have written in the past couple of posts... and thinking sometimes I feel like a duck paddling upstream in a river: There appears almost no activity on the surface and under the water the poor duck is paddling like mad to make progress against the current.

The annual report is now almost finished. I keep hoping it is finished and then there are more small changes to make. The annual report has taken a lot of my time, plus a lot of a couple of other peoples time in the UK. The overhead of red tape these days seems enormous. Gone are the days of getting on with the task and being trusted that you are getting on with the task [whatever that task is]. It's sad really - like the whole world has suddenly lost its innocence and has become a frantic fast moving bullet train.

We have also been struggling with personnel problems. One part time worker [who had been writing for us] causing us a large amount of time and effort. Since this person is part of the reason why we're here we couldn't just drop him like a lead balloon and duck and hide while the pieces fell everywhere. The fallout is still having effects on both time and energy.

In between that I have been doing some technical stuff for a new website we hoped to have up and running by October. We'll miss that deadline. The new website is 100% interactive - what people call 'Web 2.0' - very different in look and feel to anthing else we have done. The person who had been the developer on it is now on another project with us, so we have taken on a second developer to work on this. We're pretty sure we will have enough work for two developers over the next 12 months, but still this is a step into something bigger.

We have now taken on responsibility for an office in one Middle Eastern country, so that too is a step bigger. Both developers will work from this office. Its a good step and one of the plans is that the two developers will also be trained to take on the system administration for all the servers. They are starting to do this and have already relieved some of the pressure on Peter and myself. But this also means we need to train them - we did some training in August and have seen since then them taking on some of the responsibility for the system administration. Peter and I really like it when we find out they have sorted out a problem or installed something without having to come back to us for extra help information.

But... the big new Web 2.0 website does need our help and that has consumed some of my time over the last month. We heard that we should [hopefully] be getting an extra full time member of the team in January. A western trained System Administrator who will take over supervising and co-ordinating all the systems administration for us. Initially that is creating extra workload for me - getting the Job Description and other paperwork done, and trying to sort out how a visa will work for him. He will be based at out Cyprus office. We already have a desk waiting for him - a couple of weeks ago we were given some extra office furniture from another organisation here in Cyprus.

And on the change front, the office flat [used by the various Middle Eastern workers when they come here] had to be changed as the block that it was in will be pulled down this month. We have now found a new flat and moved everything to it. The new flat will be nicer - it's smaller and more compact, but much better quality. That change too took up time in my month.

My next month? Well... it will be a catch up month. Get the visa for our new worker sorted, get the web 2.0 site live, get other facilities working on our 'flagship' site, finally send in the annual report [must be by end of October] and hopefully have a slightly quieter month. Peter and I want to try to get some time for thinking/brainstorming together. Last Autumn we did this and it helped for what we did in 2008. We have a couple of things to add to it this year - one is a 'Risk Management' policy, the other a 'Reserves' policy. The reserves policy should be easy, but trying to work out risk management on what we do and how to reduce those risks, well... that's a different issue altogether!

As some of you know that I love sailing. Over the summer I have been lent an outboard motor. I used it yesterday when the wind was too light for sailing quickly back to the club. Now... if I can only find an 'outboard motor' to fix to this duck paddling desperatly upstream I'm sure we'll make more progress!