Wednesday, October 24, 2007

Perfect Storm Disk Replacement

I recently had a drive go bad in a Sun StorEdge 3510 FC JBOD array connected to a V490 running Solaris 10 with Solaris Volume Manager. The disk was part of a five-disk stripeset that was mirrored with another stripset.

It was *not* easy finding documentation for getting this done. Tools I'd used on other systems that had SCSI attached arrays and on systems with FC attached RAID arrays did not work. The combination of JBOD with FC on a 3510 managed with Solaris Volume Manager with an active hot spare made it interesting. So without further ado...

How To Replace a Failed Drive on a JBOD Sun StorEdge 3510 FC Array That Has Been Failed Over to a Hot Spare Managed by Volume Manager in Solaris 10 (whew!)

Here's the device with the bad disk c1t10d0s0 that was replaced with the hot spare from c1t11d0s0:

# metastat d15
d15: Mirror
Submirror 0: d16
State: Okay
Submirror 1: d17
State: Okay
Pass: 1
Read option: roundrobin (default)
Write option: parallel (default)
Size: 716634624 blocks (341 GB)

d16: Submirror of d15
State: Okay
Hot spare pool: hsp000
Size: 716634624 blocks (341 GB)
Stripe 0: (interlace: 256 blocks)
Device Start Block Dbase State Reloc Hot Spare
c1t4d0s0 20352 Yes Okay Yes
c1t3d0s0 20352 Yes Okay Yes
c1t2d0s0 20352 Yes Okay Yes
c1t1d0s0 20352 Yes Okay Yes
c1t0d0s0 20352 Yes Okay Yes

d17: Submirror of d15
State: Okay
Hot spare pool: hsp000
Size: 716634624 blocks (341 GB)
Stripe 0: (interlace: 256 blocks)
Device Start Block Dbase State Reloc Hot Spare
c1t9d0s0 20352 Yes Okay Yes
c1t8d0s0 20352 Yes Okay Yes
c1t7d0s0 20352 Yes Okay Yes
c1t6d0s0 20352 Yes Okay Yes
c1t10d0s0 20352 No Okay Yes c1t11d0s0

Device Relocation Information:
Device Reloc Device ID
c1t4d0 Yes id1,ssd@n20000011c6968cf9
c1t3d0 Yes id1,ssd@n20000011c6967f16
c1t2d0 Yes id1,ssd@n20000011c6968c7c
c1t1d0 Yes id1,ssd@n20000011c68baaed
c1t0d0 Yes id1,ssd@n20000011c6968ca1
c1t9d0 Yes id1,ssd@n20000011c6967e6e
c1t8d0 Yes id1,ssd@n20000011c68b0388
c1t7d0 Yes id1,ssd@n20000011c68deaaf
c1t6d0 Yes id1,ssd@n20000011c6969259
c1t11d0 Yes id1,ssd@n20000011c68bbb2d

I removed the meta database replicas that were on c1t10d0 but I'm not convinced I had to do that before continuing.

The cfgadm command can show the attachment point for the disk.

# cfgadm -al
Ap_Id Type Receptacle Occupant Condition
c0 scsi-bus connected configured unknown
c0::dsk/c0t0d0 CD-ROM connected configured unknown
c1 fc-private connected configured unknown
c1::22000011c68b0388 disk connected configured unknown
c1::22000011c68b5cb3 disk connected configured unknown
c1::22000011c68baaed disk connected configured unknown
c1::22000011c68bbb2d disk connected configured unknown
c1::22000011c68deaaf disk connected configured unknown
c1::22000011c6967e6e disk connected configured unknown
c1::22000011c6967f16 disk connected configured unknown
c1::22000011c6968c7c disk connected configured unknown
c1::22000011c6968ca1 disk connected configured unknown
c1::22000011c6968cf9 disk connected configured unknown
c1::22000011c6969259 disk connected configured unknown
c1::22000011c696a895 disk connected configured unknown
c1::225000c0ff086290 ESI connected configured unknown
c2 fc-private connected configured unknown
c2::500000e01127c191 disk connected configured unknown
c2::500000e01127c8a1 disk connected configured unknown
usb0/1 unknown empty unconfigured ok
usb0/2 unknown empty unconfigured ok
usb0/3 unknown empty unconfigured ok
usb0/4 unknown empty unconfigured ok

However, both the cfgadm and luxadm commands are unable to remove the drive since it's on a fiber loop and is a JBOD array.

# cfgadm -x replace_device c1::22000011c68b5cb3
cfgadm: Configuration operation not supported

# luxadm remove_device 22000011c68b5cb3

WARNING!!! Please ensure that no filesystems are mounted on these device(s).
All data on these devices should have been backed up.

Error: Invalid path. Device is not a SENA subsystem. - 22000011c68b5cb3.

Instead, use luxadm to offline the bad disk:

# luxadm -e offline /dev/rdsk/c1t10d0s2

Then devfsadm to remove the dev entries:

# devfsadm -Cv
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s0
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s1
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s2
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s3
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s4
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s5
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s6
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s7
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s0
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s1
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s2
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s3
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s4
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s5
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s6
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s7

The output from cfgadm now shows the device as unusable:

# cfgadm -al
Ap_Id Type Receptacle Occupant Condition
c0 scsi-bus connected configured unknown
c0::dsk/c0t0d0 CD-ROM connected configured unknown
c1 fc-private connected configured unknown
c1::22000011c68b0388 disk connected configured unknown
c1::22000011c68b5cb3 disk connected configured unusable
c1::22000011c68baaed disk connected configured unknown
c1::22000011c68bbb2d disk connected configured unknown
[...snip...]

Physically replace the device. In the 3510 JBOD array with the default boxid of zero (check the button hidden under the left plastic ear tab), the disk layout looks like this:

0 3 6 9
1 4 7 10
2 5 8 11

(0 to 11 counting down columns first then over rows)

When the disk is replaced, the devfsadm daemon should pick up the disk immediately and configure the dev entries. If not, try this to see what the problem is:

# luxadm -e port
/devices/pci@9,600000/SUNW,qlc@2/fp@0,0:devctl CONNECTED
/devices/pci@8,600000/SUNW,qlc@1/fp@0,0:devctl CONNECTED

Note: If you get a "NOT CONNECTED" error on the 3510 path, check cfgadm to see if the fiber connection is connected.

# cfgadm -al
Ap_Id Type Receptacle Occupant Condition
c0 scsi-bus connected configured unknown
c0::dsk/c0t0d0 CD-ROM connected configured unknown
c1 fc-private connected configured unknown
c1::22000011c68b0388 disk connected configured unknown
c1::22000011c68baaed disk connected configured unknown
c1::22000011c68bbb2d disk connected configured unknown
c1::22000011c68deaaf disk connected configured unknown
c1::22000011c6967e6e disk connected configured unknown
c1::22000011c6967f16 disk connected configured unknown
c1::22000011c6968c7c disk connected configured unknown
c1::22000011c6968ca1 disk connected configured unknown
c1::22000011c6968cf9 disk connected configured unknown
c1::22000011c6969259 disk connected configured unknown
c1::22000011c696a895 disk connected configured unknown
c1::225000c0ff086290 ESI connected configured unknown
c1::500000e014cb0282 disk connected configured unknown
c2 fc-private connected configured unknown
c2::500000e01127c191 disk connected configured unknown
c2::500000e01127c8a1 disk connected configured unknown
usb0/1 unknown empty unconfigured ok
usb0/2 unknown empty unconfigured ok
usb0/3 unknown empty unconfigured ok
usb0/4 unknown empty unconfigured ok

If the controller isn't there or is unconfigured try the following:

# cfgadm -c configure cx

If the drives appear with a condition set to "unusable" do the following using the pathname from the luxadm -e port command above:

# luxadm -e forcelip devices/pci@9,600000/SUNW,qlc@2/fp@0,0:devctl

Once the dev devices for the replaced drive are back in, use format to partition the new drive like the old one used to be. You can use the partition map from the hot spare as a template.

Once the drive is partitioned, add any database replicas that may have been on the original device (I should mention that I forgot to do that, so I'm not 100% sure that works), then do a metareplace to trigger the hot spare to go back to available and the replaced drive to start resyncing:

# metareplace -e d17 c1t10d0s0

Show progress with:

# metastat | grep %

Resync in progress: 73 % done

and see that the hot spare is available again with:

# metahs -i

# metahs -i
hsp000: 2 hot spares
Device Status Length Reloc
c1t11d0s0 Available 143349312 blocks Yes
c1t5d0s0 Available 143349312 blocks Yes

Device Relocation Information:
Device Reloc Device ID
c1t11d0 Yes id1,ssd@n20000011c68bbb2d
c1t5d0 Yes id1,ssd@n20000011c696a895

keywords: 3150 storedge storagetek solaris volume manager hot spare fc fiber channel jbod

Wednesday, August 8, 2007

Solaris Link Aggregation Update

The network guys set up the Cisco switch with LACP active for my two ports and the connection came right up. However, it seems that the load balancing is quite far from a 50/50 split across the interfaces that I expected to see. My research continues today. See my previous post for details of how this project started.

Tuesday, August 7, 2007

Solaris Link Aggregation

I'm setting up a server that will have quite a bit of network traffic. It's a Sun Microsystems V245 with four built in bge network interfaces. I've connected two of them and am hoping to aggregate them together to combine bandwidth into a logically bigger pipe.

Link aggregation used to be called Trunking in earlier versions of Solaris. Fortunately, I'm using a version of Solaris later than 10 1/06 which was the first version to natively support aggregation. Before that, one needed separate Sun Trunking software.

The Solaris System Administration Guide: IP Services contains the information you'll need to do this, though there are a couple of typos in the manual to work around.

Quick and dirty:

If you want to include a live network connection in the aggregate, you have to unplumb it first with "ifconfig bge0 unplumb" for example. You need to be on the console since that will drop your connection, of course.

Do "eeprom local-mac-address?" to make sure it's true. If it's not, do "eeprom local-mac-address?=true".

Your interfaces must be of the type bge, e1000g, or xge, and must run at the same speed and in full duplex mode (check with "dladm show-link").

Next, set up the aggregate interface with "dladm create-aggr -d bge0 -d bge1 1". That will set up an interface called "aggr1" with both physical interfaces, as shown with "dladm show-aggr".

Finally, do a "dladm modify-aggr -l passive 1", assuming that you'll be making the switch that you're connected to (see below) "active" for LACP. I think you can make both sides active or make the host active and the switch passive. I don't suppose both sides can be passive or no negotiation would take place.

For IPv4 addresses, create /etc/hostname.aggr1 (not /etc/hostname.aggr.1 as shown in the manual) with the hostname of the server in the file, matching the hostname to ip definition in /etc/hosts. Touch /reconfigure and reboot or "reboot -- -r" to do a reconfiguration reboot.

Do an "ifconfig -a" to show that aggr1 is now the defined interface with your correct mask.

If your interfaces are connected to a switch, as mine are, you need to configure the switch ports to be used as an aggregation, and if the switch supports LACP, if must be configured in either active or passive mode (either, but not off mode).

I configured my aggregation and the links are up and running. However, since my network guy hasn't configured the switch yet, I'm getting "WARNING: IP: Hardware address 'xx.xx.xx.xx.xx.xx' trying to be our address xxx.xxx.xxx.xxx!" messages in the messages log. They should go away when the switch is configured properly. Hopefully that won't be too much of a hassle for our guys.

Monday, July 30, 2007

IBM p660 7026-6H1 RAM Installation

We have an IBM p660 model 7026-6H1 that required 2GB more RAM. The system started out with 2GB in it. A couple years ago, we had IBM install 2GB more. At that time, IBM was still selling the new DIMMs. We bought it from them and the CE came in to do the installation.
Due to application upgrades, we needed 2GB more. IBM no longer sells new memory for that box. (By the way, have you ever seen a company retire hardware faster than IBM?) Off to the refurb market, I purchased 2GB more (4 x 512MB DIMMs FRU 0000033P3584). The DIMMs must be installed in quads.

I searched the 'net for DIY installation instructions. Ha! IBM keeps a tight lid on such documents. Sun Microsystems, for contrast, keeps an online library of every document under the sun, no pun intended, and ships out CDs with the servers with animations of how to install whatever you want to. Not IBM. No sir-ee, that there computer is far to complicated for anyone except a $300/hour IBM engineer to work on. We could show you the documents, but you'd only hurt yourself.

Anyway, I was watching when the CE installed the RAM a couple years ago, so I figured I could take a pretty good whack at it. I scheduled the downtime and grabbed my toolkit and static strap.

Our p660 is a multi-processor box. There's a special rule about those boxes with single processors and how much memory they can hold, so I can't help you there. Our system is a four-way and the CPU shelf (not the I/O shelf) contains a 16-slot memory expansion board.

I had to unplug a couple of items like the keyboard and mouse to get the CPU shelf to slide out the back of the rack far enough to get the top off. Two easy screws, no problem. The RAM expansion board is on the right, looking at it from the back. I think there are two in there, actually. I pulled up on the tabs for the leftmost one. It had sixteen slots, eight already occupied with 512MB DIMMs - two fore and aft slots in each of columns 1, 2, 7 and 8 from left to right as viewed from the rear. The memory has to be installed in sets of four, symmetrically left and right about the center line. (I know this because I installed the DIMMs symmetrically forward and back and it no worky.) I put a DIMM in the fore and aft slots in columns 3 and 6 and booted to see through "bootinfo -r" and "lsattr -El sys0 -a realmem" that the system was now showing 6GB installed.

DIY is much cheaper than purchasing an MES to have IBM do the installation. This of course doesn't address the issue of how IBM will feel when/if RAM goes bad, I call them for contract service, and they find refurbished RAMs in the box installed by "unqualified" personnel.

Gotta' Start Somewhere

I depend on Google and the Internet to do my job. I'm a systems admin and it's hard to remember how I ever did my job without the ability to search the Internet and find people having problems just like me and their solutions to those problems. I suppose I did it more slowly. In fact, I was not a system administrator when the web was born, so that's probably the reason I can't remember what it was like.

Around 1993, when I started playing with Mosaic and installed my employer's first web server and homepage, I was just the help desk guy and also helped out with some systems stuff, some network stuff, and some database programming stuff. Now, fourteen years later (yikes!) our web page has long been in the hands of others, but the web server itself is still mine. I'm now the Unix (and sometimes VMS for what we have left) guy around here.

This blog will have some of my notes from work in it. You will likely see the more interesting, difficult, or entertaining computer problems I encounter. I have found countless solutions to problems on the net, in tech forums, help pages, and blogs like this. I've always felt like I should give back a little. This blog is my modest attempt to do that.

Of course, it's actually a selfish endeavor as it will help me remember what experience I've had and will give me an easy-to-search resource of my own past solutions if the problems pop up again.