Our SAN Storage admin added a LUN from our second CLARiiON to a dual-node cluster of T5140 Solaris servers but it didn't come up multipathed. Instead, it appeared as separate devices on c2 and c3. I followed this procedure to correct:
# format
Searching for disks...
The current rpm value 0 is invalid, adjusting it to 3600
The current rpm value 0 is invalid, adjusting it to 3600
done
c2t5006016941E01B4Ed0: configured with capacity of 25.00GB
c2t5006016141E01B4Ed0: configured with capacity of 24.98GB
c3t5006016041E01B4Ed0: configured with capacity of 24.98GB
c3t5006016841E01B4Ed0: configured with capacity of 25.00GB
AVAILABLE DISK SELECTIONS:
0. c1t0d0
/pci@400/pci@0/pci@8/scsi@0/sd@0,0
1. c2t5006016941E01B4Ed0
/pci@400/pci@0/pci@c/SUNW,emlxs@0/fp@0,0/ssd@w5006016941e01b4e,0
2. c2t5006016141E01B4Ed0
/pci@400/pci@0/pci@c/SUNW,emlxs@0/fp@0,0/ssd@w5006016141e01b4e,0
3. c3t5006016041E01B4Ed0
/pci@500/pci@0/pci@9/SUNW,emlxs@0/fp@0,0/ssd@w5006016041e01b4e,0
4. c3t5006016841E01B4Ed0
/pci@500/pci@0/pci@9/SUNW,emlxs@0/fp@0,0/ssd@w5006016841e01b4e,0
5. c4t600601603F301D00B8615DAABC4BDE11d0
/scsi_vhci/ssd@g600601603f301d00b8615daabc4bde11
[...snip...]
11. c4t6006016041301D0022C173CB42BFDE11d0
/scsi_vhci/ssd@g6006016041301d0022c173cb42bfde11
# ls /dev/rdsk/*s2
/dev/rdsk/c0t0d0s2 /dev/rdsk/c4t600601603F301D00502E7B0875DBDD11d0s2
/dev/rdsk/c1t0d0s2 /dev/rdsk/c4t600601603F301D007283D06C76DBDD11d0s2
/dev/rdsk/c2t5006016141E01B4Ed0s2 <-here /dev/rdsk/c4t600601603F301D00B8615DAABC4BDE11d0s2
/dev/rdsk/c2t5006016941E01B4Ed0s2 <-here /dev/rdsk/c4t600601603F301D00EEEBA37BBC4BDE11d0s2
/dev/rdsk/c3t5006016041E01B4Ed0s2 <-here /dev/rdsk/c4t6006016041301D0022C173CB42BFDE11d0s2
/dev/rdsk/c3t5006016841E01B4Ed0s2 <-here /dev/rdsk/c4t6006016041301D002C8AD75BC443DF11d0s2
/dev/rdsk/c4t600601603F301D002C28D638BC4BDE11d0s2 /dev/rdsk/c4t6006016041301D00BEF6F73CC443DF11d0s2
/dev/rdsk/c4t600601603F301D0040A246697BDBDD11d0s2 /dev/rdsk/c4t6006016041301D00C21EE519C443DF11d0s2
/dev/rdsk/c4t600601603F301D00428B785176DBDD11d0s2
# cfgadm -al -o show_SCSI_LUN
Ap_Id Type Receptacle Occupant Condition
c2 fc-fabric connected configured unknown
c2::5006016141e01b4e,0 <-here disk connected configured unknown
c2::5006016141e05590,0 disk connected configured unknown
c2::5006016141e05590,1 disk connected configured unknown
c2::5006016141e05590,2 disk connected configured unknown
c2::5006016141e05590,3 disk connected configured unknown
c2::5006016141e05590,4 disk connected configured unknown
c2::5006016141e05590,5 disk connected configured unknown
c2::5006016141e05590,6 disk connected configured unknown
c2::5006016141e05590,7 disk connected configured unknown
c2::5006016141e05590,8 disk connected configured unknown
c2::5006016141e05590,9 disk connected configured unknown
c2::5006016141e05590,10 disk connected configured unknown
c2::5006016941e01b4e,0 <-here disk connected configured unknown
c2::5006016941e05590,0 disk connected configured unknown
c2::5006016941e05590,1 disk connected configured unknown
c2::5006016941e05590,2 disk connected configured unknown
c2::5006016941e05590,3 disk connected configured unknown
c2::5006016941e05590,4 disk connected configured unknown
c2::5006016941e05590,5 disk connected configured unknown
c2::5006016941e05590,6 disk connected configured unknown
c2::5006016941e05590,7 disk connected configured unknown
c2::5006016941e05590,8 disk connected configured unknown
c2::5006016941e05590,9 disk connected configured unknown
c2::5006016941e05590,10 disk connected configured unknown
c3 fc-fabric connected configured unknown
c3::5006016041e01b4e,0 <-here disk connected configured unknown
c3::5006016041e05590,0 disk connected configured unknown
c3::5006016041e05590,1 disk connected configured unknown
c3::5006016041e05590,2 disk connected configured unknown
c3::5006016041e05590,3 disk connected configured unknown
c3::5006016041e05590,4 disk connected configured unknown
c3::5006016041e05590,5 disk connected configured unknown
c3::5006016041e05590,6 disk connected configured unknown
c3::5006016041e05590,7 disk connected configured unknown
c3::5006016041e05590,8 disk connected configured unknown
c3::5006016041e05590,9 disk connected configured unknown
c3::5006016041e05590,10 disk connected configured unknown
c3::5006016841e01b4e,0 <-here disk connected configured unknown
c3::5006016841e05590,0 disk connected configured unknown
c3::5006016841e05590,1 disk connected configured unknown
c3::5006016841e05590,2 disk connected configured unknown
c3::5006016841e05590,3 disk connected configured unknown
c3::5006016841e05590,4 disk connected configured unknown
c3::5006016841e05590,5 disk connected configured unknown
c3::5006016841e05590,6 disk connected configured unknown
c3::5006016841e05590,7 disk connected configured unknown
c3::5006016841e05590,8 disk connected configured unknown
c3::5006016841e05590,9 disk connected configured unknown
c3::5006016841e05590,10 disk connected configured unknown
Checked the multipath status of existing drives to make sure none was using that device:
# for i in `ls *s2`; do
> mpathadm show lu $i
> done | grep -i 01b4e
Error: Logical-unit c0t0d0s2 is not found.
Error: Logical-unit c1t0d0s2 is not found.
Error: Logical-unit c2t5006016141E01B4Ed0s2 is not found.
Error: Logical-unit c2t5006016941E01B4Ed0s2 is not found.
Error: Logical-unit c3t5006016041E01B4Ed0s2 is not found.
Error: Logical-unit c3t5006016841E01B4Ed0s2 is not found.
(Those errors are to be expected when examining non-multipathed devices.)
Looks good.
# cfgadm -c unconfigure c2::5006016141e01b4e
# cfgadm -c unconfigure c2::5006016941e01b4e
# cfgadm -c unconfigure c3::5006016041e01b4e
# cfgadm -c unconfigure c3::5006016841e01b4e
# devfsadm -Cvc disk
devfsadm[13115]: verbose: removing file: /dev/dsk/c4t6006016041301D0022C173CB42BFDE11d0s7
devfsadm[13115]: verbose: symlink /dev/dsk/c4t6006016041301D0022C173CB42BFDE11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d0022c173cb42bfde11:wd
devfsadm[13115]: verbose: removing file: /dev/rdsk/c4t6006016041301D0022C173CB42BFDE11d0s7
devfsadm[13115]: verbose: symlink /dev/rdsk/c4t6006016041301D0022C173CB42BFDE11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d0022c173cb42bfde11:wd,raw
devfsadm[13115]: verbose: removing file: /dev/dsk/c4t600601603F301D00EEEBA37BBC4BDE11d0
devfsadm[13115]: verbose: symlink /dev/dsk/c4t600601603F301D00EEEBA37BBC4BDE11d0s7 -> ../../devices/scsi_vhci/ssd@g600601603f301d00eeeba37bbc4bde11:h
devfsadm[13115]: verbose: removing file: /dev/rdsk/c4t600601603F301D00EEEBA37BBC4BDE11d0
devfsadm[13115]: verbose: symlink /dev/rdsk/c4t600601603F301D00EEEBA37BBC4BDE11d0s7 -> ../../devices/scsi_vhci/ssd@g600601603f301d00eeeba37bbc4bde11:h,raw
devfsadm[13115]: verbose: removing file: /dev/dsk/c4t600601603F301D00428B785176DBDD11d0s7
devfsadm[13115]: verbose: symlink /dev/dsk/c4t600601603F301D00428B785176DBDD11d0 -> ../../devices/scsi_vhci/ssd@g600601603f301d00428b785176dbdd11:wd
devfsadm[13115]: verbose: removing file: /dev/rdsk/c4t600601603F301D00428B785176DBDD11d0s7
devfsadm[13115]: verbose: symlink /dev/rdsk/c4t600601603F301D00428B785176DBDD11d0 -> ../../devices/scsi_vhci/ssd@g600601603f301d00428b785176dbdd11:wd,raw
devfsadm[13115]: verbose: removing file: /dev/dsk/c4t600601603F301D00502E7B0875DBDD11d0
devfsadm[13115]: verbose: symlink /dev/dsk/c4t600601603F301D00502E7B0875DBDD11d0s7 -> ../../devices/scsi_vhci/ssd@g600601603f301d00502e7b0875dbdd11:h
devfsadm[13115]: verbose: removing file: /dev/rdsk/c4t600601603F301D00502E7B0875DBDD11d0
devfsadm[13115]: verbose: symlink /dev/rdsk/c4t600601603F301D00502E7B0875DBDD11d0s7 -> ../../devices/scsi_vhci/ssd@g600601603f301d00502e7b0875dbdd11:h,raw
devfsadm[13115]: verbose: removing file: /dev/dsk/c4t6006016041301D00C21EE519C443DF11d0
devfsadm[13115]: verbose: symlink /dev/dsk/c4t6006016041301D00C21EE519C443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d00c21ee519c443df11:h
devfsadm[13115]: verbose: removing file: /dev/rdsk/c4t6006016041301D00C21EE519C443DF11d0
devfsadm[13115]: verbose: symlink /dev/rdsk/c4t6006016041301D00C21EE519C443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d00c21ee519c443df11:h,raw
devfsadm[13115]: verbose: removing file: /dev/dsk/c4t6006016041301D00C21EE519C443DF11d0s7
devfsadm[13115]: verbose: symlink /dev/dsk/c4t6006016041301D00C21EE519C443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d00c21ee519c443df11:wd
devfsadm[13115]: verbose: removing file: /dev/rdsk/c4t6006016041301D00C21EE519C443DF11d0s7
devfsadm[13115]: verbose: symlink /dev/rdsk/c4t6006016041301D00C21EE519C443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d00c21ee519c443df11:wd,raw
devfsadm[13115]: verbose: removing file: /dev/dsk/c4t6006016041301D00BEF6F73CC443DF11d0
devfsadm[13115]: verbose: symlink /dev/dsk/c4t6006016041301D00BEF6F73CC443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d00bef6f73cc443df11:h
devfsadm[13115]: verbose: removing file: /dev/rdsk/c4t6006016041301D00BEF6F73CC443DF11d0
devfsadm[13115]: verbose: symlink /dev/rdsk/c4t6006016041301D00BEF6F73CC443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d00bef6f73cc443df11:h,raw
devfsadm[13115]: verbose: removing file: /dev/dsk/c4t6006016041301D00BEF6F73CC443DF11d0s7
devfsadm[13115]: verbose: symlink /dev/dsk/c4t6006016041301D00BEF6F73CC443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d00bef6f73cc443df11:wd
devfsadm[13115]: verbose: removing file: /dev/rdsk/c4t6006016041301D00BEF6F73CC443DF11d0s7
devfsadm[13115]: verbose: symlink /dev/rdsk/c4t6006016041301D00BEF6F73CC443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d00bef6f73cc443df11:wd,raw
devfsadm[13115]: verbose: removing file: /dev/dsk/c4t6006016041301D002C8AD75BC443DF11d0
devfsadm[13115]: verbose: symlink /dev/dsk/c4t6006016041301D002C8AD75BC443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d002c8ad75bc443df11:h
devfsadm[13115]: verbose: removing file: /dev/rdsk/c4t6006016041301D002C8AD75BC443DF11d0
devfsadm[13115]: verbose: symlink /dev/rdsk/c4t6006016041301D002C8AD75BC443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d002c8ad75bc443df11:h,raw
devfsadm[13115]: verbose: removing file: /dev/dsk/c4t6006016041301D002C8AD75BC443DF11d0s7
devfsadm[13115]: verbose: symlink /dev/dsk/c4t6006016041301D002C8AD75BC443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d002c8ad75bc443df11:wd
devfsadm[13115]: verbose: removing file: /dev/rdsk/c4t6006016041301D002C8AD75BC443DF11d0s7
devfsadm[13115]: verbose: symlink /dev/rdsk/c4t6006016041301D002C8AD75BC443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d002c8ad75bc443df11:wd,raw
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016041E01B4Ed0s0
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016041E01B4Ed0s1
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016041E01B4Ed0s2
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016041E01B4Ed0s3
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016041E01B4Ed0s4
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016041E01B4Ed0s5
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016041E01B4Ed0s6
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016041E01B4Ed0s7
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016841E01B4Ed0s0
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016841E01B4Ed0s1
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016841E01B4Ed0s2
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016841E01B4Ed0s3
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016841E01B4Ed0s4
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016841E01B4Ed0s5
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016841E01B4Ed0s6
devfsadm[13115]: verbose: removing file: /dev/dsk/c3t5006016841E01B4Ed0s7
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016941E01B4Ed0s0
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016941E01B4Ed0s1
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016941E01B4Ed0s2
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016941E01B4Ed0s3
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016941E01B4Ed0s4
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016941E01B4Ed0s5
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016941E01B4Ed0s6
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016941E01B4Ed0s7
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016141E01B4Ed0s0
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016141E01B4Ed0s1
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016141E01B4Ed0s2
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016141E01B4Ed0s3
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016141E01B4Ed0s4
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016141E01B4Ed0s5
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016141E01B4Ed0s6
devfsadm[13115]: verbose: removing file: /dev/dsk/c2t5006016141E01B4Ed0s7
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016041E01B4Ed0s0
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016041E01B4Ed0s1
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016041E01B4Ed0s2
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016041E01B4Ed0s3
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016041E01B4Ed0s4
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016041E01B4Ed0s5
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016041E01B4Ed0s6
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016041E01B4Ed0s7
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016841E01B4Ed0s0
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016841E01B4Ed0s1
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016841E01B4Ed0s2
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016841E01B4Ed0s3
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016841E01B4Ed0s4
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016841E01B4Ed0s5
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016841E01B4Ed0s6
devfsadm[13115]: verbose: removing file: /dev/rdsk/c3t5006016841E01B4Ed0s7
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016941E01B4Ed0s0
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016941E01B4Ed0s1
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016941E01B4Ed0s2
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016941E01B4Ed0s3
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016941E01B4Ed0s4
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016941E01B4Ed0s5
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016941E01B4Ed0s6
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016941E01B4Ed0s7
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016141E01B4Ed0s0
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016141E01B4Ed0s1
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016141E01B4Ed0s2
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016141E01B4Ed0s3
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016141E01B4Ed0s4
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016141E01B4Ed0s5
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016141E01B4Ed0s6
devfsadm[13115]: verbose: removing file: /dev/rdsk/c2t5006016141E01B4Ed0s7
# cfgadm -c configure c2::5006016141e01b4e
# cfgadm -c configure c2::5006016941e01b4e
# cfgadm -c configure c3::5006016041e01b4e
# cfgadm -c configure c3::5006016841e01b4e
# devfsadm -Cvc disk
devfsadm[13122]: verbose: removing file: /dev/dsk/c4t6006016041301D00C21EE519C443DF11d0
devfsadm[13122]: verbose: symlink /dev/dsk/c4t6006016041301D00C21EE519C443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d00c21ee519c443df11:h
devfsadm[13122]: verbose: removing file: /dev/rdsk/c4t6006016041301D00C21EE519C443DF11d0
devfsadm[13122]: verbose: symlink /dev/rdsk/c4t6006016041301D00C21EE519C443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d00c21ee519c443df11:h,raw
devfsadm[13122]: verbose: removing file: /dev/dsk/c4t6006016041301D00C21EE519C443DF11d0s7
devfsadm[13122]: verbose: symlink /dev/dsk/c4t6006016041301D00C21EE519C443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d00c21ee519c443df11:wd
devfsadm[13122]: verbose: removing file: /dev/rdsk/c4t6006016041301D00C21EE519C443DF11d0s7
devfsadm[13122]: verbose: symlink /dev/rdsk/c4t6006016041301D00C21EE519C443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d00c21ee519c443df11:wd,raw
devfsadm[13122]: verbose: removing file: /dev/dsk/c4t6006016041301D00BEF6F73CC443DF11d0
devfsadm[13122]: verbose: symlink /dev/dsk/c4t6006016041301D00BEF6F73CC443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d00bef6f73cc443df11:h
devfsadm[13122]: verbose: removing file: /dev/rdsk/c4t6006016041301D00BEF6F73CC443DF11d0
devfsadm[13122]: verbose: symlink /dev/rdsk/c4t6006016041301D00BEF6F73CC443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d00bef6f73cc443df11:h,raw
devfsadm[13122]: verbose: removing file: /dev/dsk/c4t6006016041301D00BEF6F73CC443DF11d0s7
devfsadm[13122]: verbose: symlink /dev/dsk/c4t6006016041301D00BEF6F73CC443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d00bef6f73cc443df11:wd
devfsadm[13122]: verbose: removing file: /dev/rdsk/c4t6006016041301D00BEF6F73CC443DF11d0s7
devfsadm[13122]: verbose: symlink /dev/rdsk/c4t6006016041301D00BEF6F73CC443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d00bef6f73cc443df11:wd,raw
devfsadm[13122]: verbose: removing file: /dev/dsk/c4t6006016041301D002C8AD75BC443DF11d0
devfsadm[13122]: verbose: symlink /dev/dsk/c4t6006016041301D002C8AD75BC443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d002c8ad75bc443df11:h
devfsadm[13122]: verbose: removing file: /dev/rdsk/c4t6006016041301D002C8AD75BC443DF11d0
devfsadm[13122]: verbose: symlink /dev/rdsk/c4t6006016041301D002C8AD75BC443DF11d0s7 -> ../../devices/scsi_vhci/ssd@g6006016041301d002c8ad75bc443df11:h,raw
devfsadm[13122]: verbose: removing file: /dev/dsk/c4t6006016041301D002C8AD75BC443DF11d0s7
devfsadm[13122]: verbose: symlink /dev/dsk/c4t6006016041301D002C8AD75BC443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d002c8ad75bc443df11:wd
devfsadm[13122]: verbose: removing file: /dev/rdsk/c4t6006016041301D002C8AD75BC443DF11d0s7
devfsadm[13122]: verbose: symlink /dev/rdsk/c4t6006016041301D002C8AD75BC443DF11d0 -> ../../devices/scsi_vhci/ssd@g6006016041301d002c8ad75bc443df11:wd,raw
# format
Searching for disks...done
c4t6006016008501E0016DD8AD9436ADF11d0: configured with capacity of 25.00GB
AVAILABLE DISK SELECTIONS:
0. c1t0d0
/pci@400/pci@0/pci@8/scsi@0/sd@0,0
1. c4t600601603F301D00B8615DAABC4BDE11d0
/scsi_vhci/ssd@g600601603f301d00b8615daabc4bde11
[...snip...]
11. c4t6006016041301D0022C173CB42BFDE11d0
/scsi_vhci/ssd@g6006016041301d0022c173cb42bfde11
12. c4t6006016008501E0016DD8AD9436ADF11d0
/scsi_vhci/ssd@g6006016008501e0016dd8ad9436adf11
Looking better. Though, what was all that removing of a c4... device both times with the devfsadm command?? Some device link hanging out there that doesn't want to go away? Hmmm...
Friday, May 28, 2010
Monday, May 24, 2010
Handy Solaris Hardware Troubleshooting Commands
ipmitool - utility for controlling IPMI-enabled devices
ipmitool chassis status
ipmitool fru
ipmitool pef status
ipmitool pef list
ipmitool sel info
ipmitool sel elist
ipmitool sdr list all info
ipmitool sunoem led get
ipmitool sunoem sbled get
ipmitool chassis status
ipmitool fru
ipmitool pef status
ipmitool pef list
ipmitool sel info
ipmitool sel elist
ipmitool sdr list all info
ipmitool sunoem led get
ipmitool sunoem sbled get
Friday, May 21, 2010
NFS client on OpenVMS limited to 32-bit NFSv2
I recently had a chance to dig into OpenVMS, getting back to days gone by. A project required transfer of data from one OpenVMS verion 8 host to another, and I decided to take a crack at setting up NFS between the two nodes.
It didn't take too long once I found this document from HP.
But then I got a nasty surprise when I tried to copy a 4GB file from client to server:
%COPY-E-WRITEERR, error writing DNFS2:[000000]FILENAME.DAT;1
-RMS-F-FUL, device full (insufficient space for allocation)
%COPY-W-NOTCMPLT, DRA0:[000000.TEST]FILENAME.DAT;1 not completely copied
The target drive was on a brand new server and had 140GB free, but only about half of the file was there. So what gives?
A search turned up the OpenVMS UCX TCPIP Services v5.6 release notes and this section:
3.7.2 NFS Client Problems and Restrictions
[...snip...]
* The NFS client included with TCP/IP Services uses the NFS Version 2 protocol only.
* With the NFS Version 2 protocol, the value of the file size is limited to 32 bits.
[...snip...]
With 1 of the bits to track file locking, that leaves 31 bits for the size of the file, or 2.1GB.
I expect better of OpenVMS.
It didn't take too long once I found this document from HP.
But then I got a nasty surprise when I tried to copy a 4GB file from client to server:
%COPY-E-WRITEERR, error writing DNFS2:[000000]FILENAME.DAT;1
-RMS-F-FUL, device full (insufficient space for allocation)
%COPY-W-NOTCMPLT, DRA0:[000000.TEST]FILENAME.DAT;1 not completely copied
The target drive was on a brand new server and had 140GB free, but only about half of the file was there. So what gives?
A search turned up the OpenVMS UCX TCPIP Services v5.6 release notes and this section:
3.7.2 NFS Client Problems and Restrictions
[...snip...]
* The NFS client included with TCP/IP Services uses the NFS Version 2 protocol only.
* With the NFS Version 2 protocol, the value of the file size is limited to 32 bits.
[...snip...]
With 1 of the bits to track file locking, that leaves 31 bits for the size of the file, or 2.1GB.
I expect better of OpenVMS.
Thursday, May 6, 2010
Solaris 10 FC disk/LUN Commands
Here are some commands related to diagnosing problems with FC disks/LUNs and showing multpath status if using Sun's stmsboot/MPXIO/Multipathing software:
# luxadm display /dev/rdsk/c0t0d0s2
# mpathadm show lu /dev/rdsk/c4t600...dd11d0
# mpathadm list mpath-support
# mpathadm list lu /dev/rdsk/c4t600...dd11d0
# mpathadm show mpath-support libmpscsi-vhci.so
# mpathadm show initiator-port 2101...4f93
# fcinfo hba-port
# fcinfo remote-port -l -s -p 2101...4f93
I found many useful when having to check status of and/or compare the paths.
# luxadm display /dev/rdsk/c0t0d0s2
# mpathadm show lu /dev/rdsk/c4t600...dd11d0
# mpathadm list mpath-support
# mpathadm list lu /dev/rdsk/c4t600...dd11d0
# mpathadm show mpath-support libmpscsi-vhci.so
# mpathadm show initiator-port 2101...4f93
# fcinfo hba-port
# fcinfo remote-port -l -s -p 2101...4f93
I found many useful when having to check status of and/or compare the paths.
Solaris 10 SMTP Server Not Running After Patch
About a month ago, I installed over one hundred patches on a Solaris 10 box, including patch 142436-03. Just today I discovered that the server was refusing incoming SMTP connections. (It gets incoming mail only infrequently, obviously.)
I tried to connect to port 25 from another server. No dice. I tried this from the local server:
# mconnect localhost (worked!)
# mconnect actual_hostname (didn't work)
From a Solaris 10 Discussion forum post at http://72.5.124.102/thread.jspa?threadID=5233087, I discovered that the patch must've turned on the local_only configuration by default in an effort to help out with security. I think it would be nice if a patch left things the way they were and maybe prompted you to change it. Oh well. Here's the fix:
# svccfg -s sendmail listprop | grep local_only (to verify it's set to true)
# svccfg -s sendmail setprop config/local_only = false
# svcadm refresh sendmail
# svcadm restart sendmail
Voila!
I tried to connect to port 25 from another server. No dice. I tried this from the local server:
# mconnect localhost (worked!)
# mconnect actual_hostname (didn't work)
From a Solaris 10 Discussion forum post at http://72.5.124.102/thread.jspa?threadID=5233087, I discovered that the patch must've turned on the local_only configuration by default in an effort to help out with security. I think it would be nice if a patch left things the way they were and maybe prompted you to change it. Oh well. Here's the fix:
# svccfg -s sendmail listprop | grep local_only (to verify it's set to true)
# svccfg -s sendmail setprop config/local_only = false
# svcadm refresh sendmail
# svcadm restart sendmail
Voila!
Thursday, April 29, 2010
Cleaning up Solaris 10 Device Tree when LUNs Removed
The following was copied from Symantec here: http://sfdoccentral.symantec.com/sf/5.0MP3/solaris/html/vxvm_admin/ch02s24s03.htm but I added some notes because their instructions were not correct (at least on my system) in some spots.
To clean up the device tree after you remove LUNs
1.
The removed devices show up as drive not available (or drive type unknown) in the output of the format command:
413. c3t5006048ACAFE4A7Cd252
/pci@1d,700000/SUNW,qlc@1,1/fp@0,0/ssd@w5006048acafe4a7c,fc
2.
After the LUNs are unmapped using Array management or the command line, Solaris also displays the devices as either
unusable or failing (or maybe unknown just like all the devices - make sure you have the right ones!).
bash-3.00# cfgadm -al -o show_SCSI_LUN
[...]
c2::5006048acafe4a73,256 disk connected configured unusable
c3::5006048acafe4a7c,255 disk connected configured unusable
[...]
3.
If the removed LUNs show up as failing, you need to force a LIP on the HBA. This operation probes the
targets again, so that the device shows up as unusable. Unless the device shows up as unusable, it cannot be
removed from the device tree. Do a long listing of the rdsk directory to see what device to spevify:
luxadm -e forcelip /devices/pci@1d,700000/SUNW,qlc@1,1/fp@0,0:devctl
4.
To remove the device from the cfgadm database, run the following commands on the HBA:
cfgadm -c unconfigure -o unusable_SCSI_LUN c2::5006048acafe4a73
or this one if not unusable:
cfgadm -c unconfigure -o c3::5006048acafe4a7c
5.
Repeat step 2 to verify that the LUNs have been removed.
6.
Clean up the device tree. The following command removes the /dev/rdsk... links to /devices.
$devfsadm -Cv
To clean up the device tree after you remove LUNs
1.
The removed devices show up as drive not available (or drive type unknown) in the output of the format command:
413. c3t5006048ACAFE4A7Cd252
/pci@1d,700000/SUNW,qlc@1,1/fp@0,0/ssd@w5006048acafe4a7c,fc
2.
After the LUNs are unmapped using Array management or the command line, Solaris also displays the devices as either
unusable or failing (or maybe unknown just like all the devices - make sure you have the right ones!).
bash-3.00# cfgadm -al -o show_SCSI_LUN
[...]
c2::5006048acafe4a73,256 disk connected configured unusable
c3::5006048acafe4a7c,255 disk connected configured unusable
[...]
3.
If the removed LUNs show up as failing, you need to force a LIP on the HBA. This operation probes the
targets again, so that the device shows up as unusable. Unless the device shows up as unusable, it cannot be
removed from the device tree. Do a long listing of the rdsk directory to see what device to spevify:
luxadm -e forcelip /devices/pci@1d,700000/SUNW,qlc@1,1/fp@0,0:devctl
4.
To remove the device from the cfgadm database, run the following commands on the HBA:
cfgadm -c unconfigure -o unusable_SCSI_LUN c2::5006048acafe4a73
or this one if not unusable:
cfgadm -c unconfigure -o c3::5006048acafe4a7c
5.
Repeat step 2 to verify that the LUNs have been removed.
6.
Clean up the device tree. The following command removes the /dev/rdsk... links to /devices.
$devfsadm -Cv
Friday, February 29, 2008
Working With and Cleaning Out wtmpx
% last
# This example keeps only last 500 records. You might want more on a busy system
% /usr/lib/acct/fwtmp < /var/adm/wtmpx | tail -500 | /usr/lib/acct/fwtmp -ic > /tmp/wtmpx
# Test it
% last -f /tmp/wtmpx
% cat /tmp/wtmpx > /var/adm/wtmpx
# This example keeps only last 500 records. You might want more on a busy system
% /usr/lib/acct/fwtmp < /var/adm/wtmpx | tail -500 | /usr/lib/acct/fwtmp -ic > /tmp/wtmpx
# Test it
% last -f /tmp/wtmpx
% cat /tmp/wtmpx > /var/adm/wtmpx
Thursday, February 28, 2008
Weekend Down in Flames Revisited
The error from the previous post came back, and all four drives went bad again. This time though, the field engineer replaced the I/O expansion boards and the cable connecting them.
When I went to reboot, it went to book off of dkc0. You might remember that last time, dkc0 had gone bad during the field engineer's fiddling and I had booted off of dkc1 and used volrootmir to remirror dkc0.
Well, the boot didn't go so well. It said it found a valid boot block, but when LSM went to load, it spit out errors about a bad boot track and unmirrored something or other and then went into single user mode. Already running 10 minutes late getting the system back online, I just booted from dkc1 again and it worked fine.
I then removed rz16 (dkc0) from the LSM mirror and then readded it and remirrored. The volrootmir command went something like this, this time:
# volrootmir -a rz16
INFO: The '-a' option was specified for the /usr/sbin/volrootmir command,
however there are partitions on dsk1 that are not
encapsulated and therefore can not be mirrored.
INFO: The '-a' option was specified for the /usr/sbin/volrootmir command,
however the volume on dsk11 is not in the rootdg disk
group and will not be mirrored.
INFO: The '-a' option was specified for the /usr/sbin/volrootmir command,
however the volume on dsk14 is not in the rootdg disk
group and will not be mirrored.
INFO: The '-a' option was specified for the /usr/sbin/volrootmir command,
however the volume on dsk15 is not in the rootdg disk
group and will not be mirrored.
INFO: The '-a' option was specified for the /usr/sbin/volrootmir command,
however the volume on dsk10 is not in the rootdg disk
group and will not be mirrored.
Mirroring system disk dsk1 to disk rz16.
Mirroring rootvol to rz16a.
Mirroring swapvol to rz16b.
Mirroring vol-rz16g to rz16g.
Hmmm. It appears to have worked fine, though there's that note about partitions not being encapsulated. I assume it's referring to the empty partitions or the LSMsimp partition.
I'm left with a sinking feeling though. If the dkc1 disk I booted from goes bad, will I be able to boot from dkc0, or will I face a world of hurt and effort trying to boot from CD and restore from tape? I'd better schedule some downtime soon to try and boot from dkc0 and file a call with HP if it doesn't work. Just four more months with this server before we retire it!
In the meantime, the users are sucking up disk space faster than a "Stand by Me" leech sucks balls. I grabbed a couple of the unused 4.3GB drives, added them to the LSM config in voldiskadm as prod09 and prod10, then did:
# volassist make prodvol-09 8373900s prod09
# volassist -g prod mirror prodvol-09 prod10
# addvol /dev/vol/prod/prodvol-09 gfs_prod
I have about 8GB more disk space left. I hope it's enough to last four months.
When I went to reboot, it went to book off of dkc0. You might remember that last time, dkc0 had gone bad during the field engineer's fiddling and I had booted off of dkc1 and used volrootmir to remirror dkc0.
Well, the boot didn't go so well. It said it found a valid boot block, but when LSM went to load, it spit out errors about a bad boot track and unmirrored something or other and then went into single user mode. Already running 10 minutes late getting the system back online, I just booted from dkc1 again and it worked fine.
I then removed rz16 (dkc0) from the LSM mirror and then readded it and remirrored. The volrootmir command went something like this, this time:
# volrootmir -a rz16
INFO: The '-a' option was specified for the /usr/sbin/volrootmir command,
however there are partitions on dsk1 that are not
encapsulated and therefore can not be mirrored.
INFO: The '-a' option was specified for the /usr/sbin/volrootmir command,
however the volume on dsk11 is not in the rootdg disk
group and will not be mirrored.
INFO: The '-a' option was specified for the /usr/sbin/volrootmir command,
however the volume on dsk14 is not in the rootdg disk
group and will not be mirrored.
INFO: The '-a' option was specified for the /usr/sbin/volrootmir command,
however the volume on dsk15 is not in the rootdg disk
group and will not be mirrored.
INFO: The '-a' option was specified for the /usr/sbin/volrootmir command,
however the volume on dsk10 is not in the rootdg disk
group and will not be mirrored.
Mirroring system disk dsk1 to disk rz16.
Mirroring rootvol to rz16a.
Mirroring swapvol to rz16b.
Mirroring vol-rz16g to rz16g.
Hmmm. It appears to have worked fine, though there's that note about partitions not being encapsulated. I assume it's referring to the empty partitions or the LSMsimp partition.
I'm left with a sinking feeling though. If the dkc1 disk I booted from goes bad, will I be able to boot from dkc0, or will I face a world of hurt and effort trying to boot from CD and restore from tape? I'd better schedule some downtime soon to try and boot from dkc0 and file a call with HP if it doesn't work. Just four more months with this server before we retire it!
In the meantime, the users are sucking up disk space faster than a "Stand by Me" leech sucks balls. I grabbed a couple of the unused 4.3GB drives, added them to the LSM config in voldiskadm as prod09 and prod10, then did:
# volassist make prodvol-09 8373900s prod09
# volassist -g prod mirror prodvol-09 prod10
# addvol /dev/vol/prod/prodvol-09 gfs_prod
I have about 8GB more disk space left. I hope it's enough to last four months.
Monday, January 7, 2008
Eudora no longer <<Dominant>>
Late last week, I began to get this error from Eudora on my Mac:
"error while performing unknown task for <<Dominant>>"
My sponsored version of Eudora was trying every two minutes to contact adserver.eudora.com to fetch an updated advertisement. The problem is, Qualcomm no longer develops nor supports Eudora, and apparently turned off the adserver server. Without a response from the server, the Eudora client would pop up an error every couple minutes.
As a quick-n-dirty patch, I put an entry in my Mac's /etc/hosts file for "adserver.eudora.com" with an IP number that belongs to one of our local web servers and restarted Eudora. I haven't received the error since then. The access log on the web server shows that my Mac is requesting:
"POST /adjoin/playlists HTTP/1.0" 404 285
The web server is obviously replying with a 404 error, but at least Eudora doesn't timeout now. Not the best fix, but the best for now.
"error while performing unknown task for <<Dominant>>"
My sponsored version of Eudora was trying every two minutes to contact adserver.eudora.com to fetch an updated advertisement. The problem is, Qualcomm no longer develops nor supports Eudora, and apparently turned off the adserver server. Without a response from the server, the Eudora client would pop up an error every couple minutes.
As a quick-n-dirty patch, I put an entry in my Mac's /etc/hosts file for "adserver.eudora.com" with an IP number that belongs to one of our local web servers and restarted Eudora. I haven't received the error since then. The access log on the web server shows that my Mac is requesting:
"POST /adjoin/playlists HTTP/1.0" 404 285
The web server is obviously replying with a 404 error, but at least Eudora doesn't timeout now. Not the best fix, but the best for now.
Wednesday, January 2, 2008
Back to that Pesky Virtual Frame Buffer
Sigh. While my procedures with the Solaris virtual frame buffer (see previous posts) has usually been working okay, I had to modify the kill command to make it definitely exclude the grep process even though it should have excluded it anyway - very weird:
/usr/bin/kill `ps -ef | grep Xsun | grep -v grep | grep ":5" | awk '{print $2}'`
The user reports that they're occasionally getting this error in their logs:
fstat: Bad file number
(failed to stat vfb)
stat: No such file or directory
(failed to stat vfb)
fstat: Bad file number
(failed to stat vfb)
stat: No such file or directory
(failed to stat vfb)
X connection to localhost:5.0 broken (explicit kill or server shutdown).
This is due to the VFB running in a local Solaris zone. According to a post I found at http://forum.java.sun.com/thread.jspa?threadID=5233796&tstart=120 one can add /dev/winlock to the local zone this way to solve those errors:
"Steps to add the pseudo device '/dev/winlock' to the local zone:
As a superuser in the global zone,
1. Add the '/dev/winlock' pseudo device to the local zone:
global# zonecfg -z
zonecfg:zonename> add device
zonecfg:zonename:device> set match=/dev/winlock
zonecfg:zonename:device> end
zonecfg:zonename> exit
With this, the local zone will have access to '/dev/winlock' device.
2. Reboot the local zone
global# zoneadm -z reboot
Once the local zone is rebooted, '/dev/winlock' will now be available in the local zone."
I've made the change in the zone configuration so it will take effect the next time I reboot. I can see it now - I'll reboot, the instructions above will have been wrong, the zone won't come up, and it'll take me a half hour to remember what changes I made... hopefully I've made enough notes that I'll remember it.
/usr/bin/kill `ps -ef | grep Xsun | grep -v grep | grep ":5" | awk '{print $2}'`
The user reports that they're occasionally getting this error in their logs:
fstat: Bad file number
(failed to stat vfb)
stat: No such file or directory
(failed to stat vfb)
fstat: Bad file number
(failed to stat vfb)
stat: No such file or directory
(failed to stat vfb)
X connection to localhost:5.0 broken (explicit kill or server shutdown).
This is due to the VFB running in a local Solaris zone. According to a post I found at http://forum.java.sun.com/thread.jspa?threadID=5233796&tstart=120 one can add /dev/winlock to the local zone this way to solve those errors:
"Steps to add the pseudo device '/dev/winlock' to the local zone:
As a superuser in the global zone,
1. Add the '/dev/winlock' pseudo device to the local zone:
global# zonecfg -z
zonecfg:zonename> add device
zonecfg:zonename:device> set match=/dev/winlock
zonecfg:zonename:device> end
zonecfg:zonename> exit
With this, the local zone will have access to '/dev/winlock' device.
2. Reboot the local zone
global# zoneadm -z
Once the local zone is rebooted, '/dev/winlock' will now be available in the local zone."
I've made the change in the zone configuration so it will take effect the next time I reboot. I can see it now - I'll reboot, the instructions above will have been wrong, the zone won't come up, and it'll take me a half hour to remember what changes I made... hopefully I've made enough notes that I'll remember it.
Monday, December 31, 2007
Weekend Down in Flames
A page at 2:46 Saturday morning woke me up. Our Lawson server, an aging Digital Alpha 4100 running Tru64 Unix 5.1B was complaining. Right away I knew that my list of chores I had planned for the weekend was going to take a hit.
A volume managed in LSM (Logical Storage Manager) had lost a plex and one of the secondary swap spaces was running unmirrored. It looked like a relatively simple disk failure, so I headed back to bed.
In the morning, I went to the office to gather up some paperwork, then drove to the computer room. The alarm in the disk cabinet was screaming like a banshee, and it was festooned with a number of amber lights. The cabinet contains two sets of redundant HSZ70 scsi controllers. The top set is connected to an expansion unit on the back of the cabinet. There are about 70 disks in the cabinet. Four of them, all on channel 1 off the expansion unit were blinking amber. Yikes!
I checked the CLI interface (Command Line Interface Interface) and could see that two raidsets were running reduced, the third disk was the one that was the swap mirror, and the fourth bad disk isn't used. Since everything was mirrored or raided, the system was still up and functioning. I filed a call with HP support at 1-800-354-9000 (1-800-DIAL-A-PRAYER). I tried for a while to convince the phone support tech that there was probably something wrong with the entire channel on the expansion unit, but the controller didn't have any errors on it, so he eventually sent the call to the local field tech and asked them to bring over four new drives. Uh huh. Fortunately, I know the field tech is better than that, and when he called, I explained the situation.
Later in the evening, the tech arrived at the computer room (I had since gone back home) and he unplugged the expansion unit, blew on it, and put it back in. That brought the disks back online. However, it also managed to take down all the disks on channel one on the front of the cabinet too, including the mirror of the root disk that I'd booted from. He rebuilt all the disks from the hardware side, but it left a mighty mess to clean up on the software side.
With the boot disk down and "replaced", that means a shutdown and boot from the other root mirror to free it up so it can be rebuilt.
1) Disassociate and remove the bad plexes from the root volume
volplex -o rm dis rootvol-01
volplex -o rm dis swapvol-01
volplex -o rm dis vol-rz16g-01
2) Remove the disks from the diskgroup
voldg rmdisk rz16a
voldg rmdisk rz16b
voldg rmdisk rz16f
voldg rmdisk rz16g
3) Remove disks from LSM:
voldisk rm rz16a
voldisk rm rz16b
voldisk rm rz16f
voldisk rm rz16g
4) Physically replace the disks using the HSZ commands (I didn't have to replace mine since it wasn't the disks that failed)
5) Label the new disk
disklabel -wr rz16
If that doesn't work, try:
disklabel -z rz16
disklabel -wr rz16
For me, trying to label the disk didn't work at all because it was the boot disk and the OS still had a partition open and claimed. I had to "shutdown -h now" and reboot from the >>> console prompt using the address of the mirrored (good) root drive. On my system a show bootdef_dev showed that the default boot disk was dkc0..., so I used "boot dkc1" to boot from the mirror.
When the system rebooted, I was then able to continue with the disklabel:
disklabel -wr rz16
6) Mirror the root drive
volrootmir -a rz16
That command will build a mirror on the disk specified.
After all that, I tried to remirror the additional swap volume, but ran into a roadblock.
The disks were set up in LSM when the system was installed using Tru64 4.0x. I had since upgraded to Tru64 5.1B. LSM used to use a private region of 1024, but the new version uses a private region of 4096. When I tried to add the new disk using voldiskadm, and then mirror it, it said there wasn't enough space to complete the mirror:
# volassist mirror swapvol02 swap01
lsm:volassist: ERROR: Cannot allocate space to mirror 8376988 block volume
I then tried to write a disklabel from the good mirror onto the disk before adding it with voldiskadm (dsk13 is the good disk, dsk12 is the new one):
disklabel -r dsk13 > dsk13.lab
disklabel -R dsk12 dsk13.lab
I then added the disk through voldiskadm and chose not to initialize it. But then I got this error:
# volassist mirror swapvol02 swap01
lsm:volplex: ERROR: Volume swapvol02, plex swapvol02-01, block 0: Plex write:
Error: Write failure
lsm:volplex: ERROR: sd swap01-01 in plex swapvol02-01 failed during attach
lsm:volplex: ERROR: changing plex swapvol02-01:
Record not in disk group
lsm:volplex: ERROR: Attempting to cleanup after failure ...
lsm:volassist: ERROR: Could not attach new mirror(s) to volume swapvol02
lsm:volassist: WARNING: Object swapvol02-01: Unexpectedly removed from the configuration
Ouch. I turned back to Google and found two tech forum posts from people who'd encountered the same thing, but no one had answered them.
I called HP support and they sent me these three simple little commands:
# voldisksetup -i dsk12 privlen=1024
# voldg adddisk swap01=dsk12
# volassist mirror swapvol02 swap01
Worked like a charm, and the system is back to normal with everything mirrored and raided properly.
However, there's still an amber warning light on the cabinet. Since the controllers aren't reporting any errors, the field tech's best guess is a problem with one of the many fans in the cabinet. He'll be coming by Wednesday with a few new fans to try some replacements to see if he can get that pesky amber light to go away.
A volume managed in LSM (Logical Storage Manager) had lost a plex and one of the secondary swap spaces was running unmirrored. It looked like a relatively simple disk failure, so I headed back to bed.
In the morning, I went to the office to gather up some paperwork, then drove to the computer room. The alarm in the disk cabinet was screaming like a banshee, and it was festooned with a number of amber lights. The cabinet contains two sets of redundant HSZ70 scsi controllers. The top set is connected to an expansion unit on the back of the cabinet. There are about 70 disks in the cabinet. Four of them, all on channel 1 off the expansion unit were blinking amber. Yikes!
I checked the CLI interface (Command Line Interface Interface) and could see that two raidsets were running reduced, the third disk was the one that was the swap mirror, and the fourth bad disk isn't used. Since everything was mirrored or raided, the system was still up and functioning. I filed a call with HP support at 1-800-354-9000 (1-800-DIAL-A-PRAYER). I tried for a while to convince the phone support tech that there was probably something wrong with the entire channel on the expansion unit, but the controller didn't have any errors on it, so he eventually sent the call to the local field tech and asked them to bring over four new drives. Uh huh. Fortunately, I know the field tech is better than that, and when he called, I explained the situation.
Later in the evening, the tech arrived at the computer room (I had since gone back home) and he unplugged the expansion unit, blew on it, and put it back in. That brought the disks back online. However, it also managed to take down all the disks on channel one on the front of the cabinet too, including the mirror of the root disk that I'd booted from. He rebuilt all the disks from the hardware side, but it left a mighty mess to clean up on the software side.
With the boot disk down and "replaced", that means a shutdown and boot from the other root mirror to free it up so it can be rebuilt.
1) Disassociate and remove the bad plexes from the root volume
volplex -o rm dis rootvol-01
volplex -o rm dis swapvol-01
volplex -o rm dis vol-rz16g-01
2) Remove the disks from the diskgroup
voldg rmdisk rz16a
voldg rmdisk rz16b
voldg rmdisk rz16f
voldg rmdisk rz16g
3) Remove disks from LSM:
voldisk rm rz16a
voldisk rm rz16b
voldisk rm rz16f
voldisk rm rz16g
4) Physically replace the disks using the HSZ commands (I didn't have to replace mine since it wasn't the disks that failed)
5) Label the new disk
disklabel -wr rz16
If that doesn't work, try:
disklabel -z rz16
disklabel -wr rz16
For me, trying to label the disk didn't work at all because it was the boot disk and the OS still had a partition open and claimed. I had to "shutdown -h now" and reboot from the >>> console prompt using the address of the mirrored (good) root drive. On my system a show bootdef_dev showed that the default boot disk was dkc0..., so I used "boot dkc1" to boot from the mirror.
When the system rebooted, I was then able to continue with the disklabel:
disklabel -wr rz16
6) Mirror the root drive
volrootmir -a rz16
That command will build a mirror on the disk specified.
After all that, I tried to remirror the additional swap volume, but ran into a roadblock.
The disks were set up in LSM when the system was installed using Tru64 4.0x. I had since upgraded to Tru64 5.1B. LSM used to use a private region of 1024, but the new version uses a private region of 4096. When I tried to add the new disk using voldiskadm, and then mirror it, it said there wasn't enough space to complete the mirror:
# volassist mirror swapvol02 swap01
lsm:volassist: ERROR: Cannot allocate space to mirror 8376988 block volume
I then tried to write a disklabel from the good mirror onto the disk before adding it with voldiskadm (dsk13 is the good disk, dsk12 is the new one):
disklabel -r dsk13 > dsk13.lab
disklabel -R dsk12 dsk13.lab
I then added the disk through voldiskadm and chose not to initialize it. But then I got this error:
# volassist mirror swapvol02 swap01
lsm:volplex: ERROR: Volume swapvol02, plex swapvol02-01, block 0: Plex write:
Error: Write failure
lsm:volplex: ERROR: sd swap01-01 in plex swapvol02-01 failed during attach
lsm:volplex: ERROR: changing plex swapvol02-01:
Record not in disk group
lsm:volplex: ERROR: Attempting to cleanup after failure ...
lsm:volassist: ERROR: Could not attach new mirror(s) to volume swapvol02
lsm:volassist: WARNING: Object swapvol02-01: Unexpectedly removed from the configuration
Ouch. I turned back to Google and found two tech forum posts from people who'd encountered the same thing, but no one had answered them.
I called HP support and they sent me these three simple little commands:
# voldisksetup -i dsk12 privlen=1024
# voldg adddisk swap01=dsk12
# volassist mirror swapvol02 swap01
Worked like a charm, and the system is back to normal with everything mirrored and raided properly.
However, there's still an amber warning light on the cabinet. Since the controllers aren't reporting any errors, the field tech's best guess is a problem with one of the many fans in the cabinet. He'll be coming by Wednesday with a few new fans to try some replacements to see if he can get that pesky amber light to go away.
Wednesday, December 19, 2007
Quote of the Day
Here's my favorite quote today, found while browsing to try to find an AIX hardware support matrix. Apparently IBM is stingy with that information too.
From "inferno" on www.pseriestech.org/forum/aix:
"I do not get it, I explained that I was interested in getting IBM certified. Companies like Sun Microsystems and Red Hat will constantly harrass you and damn near send a limo and a bunch of exotic dancers to your home in order to get you certified. So what is wrong with IBM?"
From "inferno" on www.pseriestech.org/forum/aix:
"I do not get it, I explained that I was interested in getting IBM certified. Companies like Sun Microsystems and Red Hat will constantly harrass you and damn near send a limo and a bunch of exotic dancers to your home in order to get you certified. So what is wrong with IBM?"
Thursday, December 13, 2007
X11 Forwarding over SSH
Holy crap. I just spent about five hours banging my head against a wall.
In an effort to try to secure connections, I've been trying to get more things tunneled through ssh to lock down some more ports to our DMZ network. Today, I've been working on getting X Windows applications to tunnel over ssh to my PC.
I read a couple online manuals.
I connected from PuTTY on my PC to a Solaris 9 box on our DMZ. Ssh session came right up. I turned on X11 forwarding and enabled it on the server and tried to log in again. No dice. It closed the connection right after I typed in the password. I must be doing something wrong.
I did some more web searching. Read several more tutorials about ssh and X11 forwarding. Still no dice. Still must be doing something wrong. Click this. Click that. Edit this config file, edit that. Nope. Passive, active, indirect. Nope. Port forwarding. No port forwarding. Nope. DISPLAY set. DISPLAY not set. Nope. Nope. Nothing in the PuTTY logs.
I really must not understand this ssh/X11 forwarding thing at all. Yet every document I read has virtually the same instructions. What could I be missing?
I finally happen to check the error log on the Solaris 9 box itself. Sure enough, there were errors that corresponded to each of my login attempts. I did some more web searching and finally found it: a bug report for a Solaris 9 patch that causes X11 forwarding to fail. Effin' A.
I tried one of my Solaris 10 servers. Worked the first time. Five hours gone up in smoke. No wonder I'm quiet at the dinner table. I just worked hard all day doing nothing.
In an effort to try to secure connections, I've been trying to get more things tunneled through ssh to lock down some more ports to our DMZ network. Today, I've been working on getting X Windows applications to tunnel over ssh to my PC.
I read a couple online manuals.
I connected from PuTTY on my PC to a Solaris 9 box on our DMZ. Ssh session came right up. I turned on X11 forwarding and enabled it on the server and tried to log in again. No dice. It closed the connection right after I typed in the password. I must be doing something wrong.
I did some more web searching. Read several more tutorials about ssh and X11 forwarding. Still no dice. Still must be doing something wrong. Click this. Click that. Edit this config file, edit that. Nope. Passive, active, indirect. Nope. Port forwarding. No port forwarding. Nope. DISPLAY set. DISPLAY not set. Nope. Nope. Nothing in the PuTTY logs.
I really must not understand this ssh/X11 forwarding thing at all. Yet every document I read has virtually the same instructions. What could I be missing?
I finally happen to check the error log on the Solaris 9 box itself. Sure enough, there were errors that corresponded to each of my login attempts. I did some more web searching and finally found it: a bug report for a Solaris 9 patch that causes X11 forwarding to fail. Effin' A.
I tried one of my Solaris 10 servers. Worked the first time. Five hours gone up in smoke. No wonder I'm quiet at the dinner table. I just worked hard all day doing nothing.
Friday, December 7, 2007
Reason #682 Why I Get Frustrated with IBM
An application developer down the hall came in the office late yesterday asking if I could install the HP LaserJet 4000 printer drivers on one of our IBM AIX servers. "Sure, no problem," I said, knowing in my core I was about to embark on a painful adventure.
This morning, I set about trying to locate the printer drivers on IBM's website. Fifteen minutes of thrashing about... no deal. They link to a huge list of some Infoprint driver crap, but nothing for HP printers.
I checked out HP's website, which is almost as poorly designed as IBM's. IBM's is worse because not only is the design bad, but they don't let you download anything useful. HP has lots of useful stuff on their website, but it's just really hard to find.
Anyway, after a few minutes of thrashing on HP's website, I found they have some generic un*x drivers for the LaserJet 4000 series, but I knew that wasn't going to fly on the AIX server. I needed the *.rte files from IBM.
I turned to Google, and found many posts to technical forums, and each one went like this:
Question: Where on the web can I find HP printer drivers for IBM?
Answer: You can't. I think they're on the installation CDs somewhere.
I then dove into my storage cabinet and pulled out my big box o' AIX cds. I located the AIX 5.2 install cds (seven of them) for that server as well as numerous other randomly labeled media. I popped the first CD into my Mac and started searching for something that looked like printer drivers. About this time, the developer guy popped in and said that his vendor said it was on disc one. Hmmm... okay.
Longer story shorter, it's not. It's disc three. I'll say it again so try to make sure Google picks it up for you out there searching the web. The HP LaserJet 4000 series printer drivers for AIX 5.2 are on installation cd number 3 of 7.
Just pop that cd into your server, or remotely mount it via NFS like I did, and run:
# smitty printers
Printer/Plotter Devices
Install Additional Printer/Plotter Software
choose /cdrom or whatever mount point you used
hit list to choose the printer drivers you want
The trick is finding them. Once you find them, it's super easy.
This morning, I set about trying to locate the printer drivers on IBM's website. Fifteen minutes of thrashing about... no deal. They link to a huge list of some Infoprint driver crap, but nothing for HP printers.
I checked out HP's website, which is almost as poorly designed as IBM's. IBM's is worse because not only is the design bad, but they don't let you download anything useful. HP has lots of useful stuff on their website, but it's just really hard to find.
Anyway, after a few minutes of thrashing on HP's website, I found they have some generic un*x drivers for the LaserJet 4000 series, but I knew that wasn't going to fly on the AIX server. I needed the *.rte files from IBM.
I turned to Google, and found many posts to technical forums, and each one went like this:
Question: Where on the web can I find HP printer drivers for IBM?
Answer: You can't. I think they're on the installation CDs somewhere.
I then dove into my storage cabinet and pulled out my big box o' AIX cds. I located the AIX 5.2 install cds (seven of them) for that server as well as numerous other randomly labeled media. I popped the first CD into my Mac and started searching for something that looked like printer drivers. About this time, the developer guy popped in and said that his vendor said it was on disc one. Hmmm... okay.
Longer story shorter, it's not. It's disc three. I'll say it again so try to make sure Google picks it up for you out there searching the web. The HP LaserJet 4000 series printer drivers for AIX 5.2 are on installation cd number 3 of 7.
Just pop that cd into your server, or remotely mount it via NFS like I did, and run:
# smitty printers
Printer/Plotter Devices
Install Additional Printer/Plotter Software
choose /cdrom or whatever mount point you used
hit list to choose the printer drivers you want
The trick is finding them. Once you find them, it's super easy.
Thursday, December 6, 2007
Update: Virtual Frame Buffer for use with Oracle Reports Server
Well, it turns out that the handy VFB I described a couple posts down doesn't work with SQR. I found a Hyperion document that had corrected resolution settings for the virtual buffer:
/usr/openwin/bin/Xvfb :5 -dev vfb screen 0 1152x900x8 &
/usr/openwin/bin/twm -display :5 -v &
DISPLAY=:5.0; export DISPLAY
Apparently, SQR is picky about the resolution. 1152x900x8 works, 1600x1200x32 did not.
/usr/openwin/bin/Xvfb :5 -dev vfb screen 0 1152x900x8 &
/usr/openwin/bin/twm -display :5 -v &
DISPLAY=:5.0; export DISPLAY
Apparently, SQR is picky about the resolution. 1152x900x8 works, 1600x1200x32 did not.
Wednesday, November 21, 2007
X Server Basics
We have been using a commercial X Windows server on our PCs to get GUI access to our unix boxes for years. Recently though, a new set of people here need to get access to a new AIX system we have. They didn't want to spring for commercial licenses, so I introduced them to the Cygwin/X server software from http://x.cygwin.com, available at no cost under a modified GNU license.
Download the software, making sure to choose to install the inetutils and xorg-x11 portions. I also installed the openssh piece so I could use SSH to connect to servers if I wanted to.
Once installed, launch Cygwin and it'll give you a unix-like terminal interface to your PC files.
Make sure dtlogin is running on the remote unix host. If it isn't, run this on the host:
# /usr/dt/bin/dtlogin &
Then on your PC in the Cygwin window, run:
Xwin -query <remote_hostname> -from <my_pc_hostname_or_ip>
That will launch the nice GUI Xwindows login for the remote host.
Download the software, making sure to choose to install the inetutils and xorg-x11 portions. I also installed the openssh piece so I could use SSH to connect to servers if I wanted to.
Once installed, launch Cygwin and it'll give you a unix-like terminal interface to your PC files.
Make sure dtlogin is running on the remote unix host. If it isn't, run this on the host:
# /usr/dt/bin/dtlogin &
Then on your PC in the Cygwin window, run:
Xwin -query <remote_hostname> -from <my_pc_hostname_or_ip>
That will launch the nice GUI Xwindows login for the remote host.
Friday, November 16, 2007
Virtual Frame Buffer for use with Oracle Reports Server
Our DBA set up Oracle Reports Server on one of my Solaris unix servers so that our applications folks can make pretty graphs and send them out to administration. Oracle Reports Server requires a connection to an X-Server to draw the pretty graphs, even though it's running in batch. Wonderful. There's no graphics console on the server, so...
Our DBA went over to a unix workstation he runs testing on, logged in, set xhost + to allow the process (well, everyone really) to connect, and then set the DISPLAY variable in the application script to connect to the workstation for the X-Server access. Pretty kludgy solution, but it worked.
All that worked fine until someone stepped on the switch on the power strip for the workstation and it was down over the weekend without anyone noticing. I started it back up, but didn't log in and had no idea about the DISPLAY setting on the other production server. Fast forward about a week, imagine applications developers running around screaming about their graphing not working, and you've got a good picture.
Our DBA eventually remembered the DISPLAY connection he'd set up, and we got the workstation logged back in. Then I went to work finding an alternative.
I located these documents:
http://www.sun.com/bigadmin/content/submitted/virtual_buffer.html
http://www.idevelopment.info/data/Unix/General_UNIX/GENERAL_XvfbWithOracle9iAS.shtml
They were helpful, but of course, we slightly incorrect. Based on their recommendations, with a tweak to the Xvfb command to get the syntax correct, I came up with the following:
In the script that does the graphics, these three lines start up the virtual frame buffer, twm, and set up the DISPLAY variable correctly:
/usr/openwin/bin/Xvfb :5 -dev vfb screen 0 1600x1200x32 &
/usr/openwin/bin/twm -display :5 -v &
DISPLAY=:5.0; export DISPLAY
At the end of the script, this line kills the vfb to make sure it's not hanging around doing nothing:
/usr/bin/kill `ps -ef | grep Xsun | grep :5 | awk '{print $2}'` > /dev/null 2>&1
Starting up the virtual frame buffer gives the Oracle Reports Server process something to connect to, and it runs on the local host even though there's no graphics console. Nice!
Now I'll go log out of that workstation...
Our DBA went over to a unix workstation he runs testing on, logged in, set xhost + to allow the process (well, everyone really) to connect, and then set the DISPLAY variable in the application script to connect to the workstation for the X-Server access. Pretty kludgy solution, but it worked.
All that worked fine until someone stepped on the switch on the power strip for the workstation and it was down over the weekend without anyone noticing. I started it back up, but didn't log in and had no idea about the DISPLAY setting on the other production server. Fast forward about a week, imagine applications developers running around screaming about their graphing not working, and you've got a good picture.
Our DBA eventually remembered the DISPLAY connection he'd set up, and we got the workstation logged back in. Then I went to work finding an alternative.
I located these documents:
http://www.sun.com/bigadmin/content/submitted/virtual_buffer.html
http://www.idevelopment.info/data/Unix/General_UNIX/GENERAL_XvfbWithOracle9iAS.shtml
They were helpful, but of course, we slightly incorrect. Based on their recommendations, with a tweak to the Xvfb command to get the syntax correct, I came up with the following:
In the script that does the graphics, these three lines start up the virtual frame buffer, twm, and set up the DISPLAY variable correctly:
/usr/openwin/bin/Xvfb :5 -dev vfb screen 0 1600x1200x32 &
/usr/openwin/bin/twm -display :5 -v &
DISPLAY=:5.0; export DISPLAY
At the end of the script, this line kills the vfb to make sure it's not hanging around doing nothing:
/usr/bin/kill `ps -ef | grep Xsun | grep :5 | awk '{print $2}'` > /dev/null 2>&1
Starting up the virtual frame buffer gives the Oracle Reports Server process something to connect to, and it runs on the local host even though there's no graphics console. Nice!
Now I'll go log out of that workstation...
Friday, November 9, 2007
Customize Mailman Messages
To create a customized welcome message in Mailman 2.1.5 that's sent to users when they subscribe, follow these steps:
Create a directory mailman/data/lists/yourlist/en (for English language) and copy subscribeack.txt from /usr/local/mailman/templates into that directory. Customize it to whatever you like. Mailman will use this custom template for the welcome message.
Schweet. This is particularly handy for distribution-only lists where the "To post to this list..." instructions in the welcome message are confusing since regular subscribers would get rejected were they to follow those instructions.
Create a directory mailman/data/lists/yourlist/en (for English language) and copy subscribeack.txt from /usr/local/mailman/templates into that directory. Customize it to whatever you like. Mailman will use this custom template for the welcome message.
Schweet. This is particularly handy for distribution-only lists where the "To post to this list..." instructions in the welcome message are confusing since regular subscribers would get rejected were they to follow those instructions.
Purge Those MySQL Binary Logs
I'm sure this is old hat to real MySQL people, but I'm pretty new to MySQL, especially replication, and our web server is usually pretty quiet, so I was a little surprised when I got a disk space warning because the binary logs had grown so large.
Signing onto the slave server, I ran "show slave status;" at the mysql> prompt to show that the server was reading from the binary log called "mysql-bin.004" on the master.
Logging onto the master, I ran "show master logs;" at the mysql> prompt (show binary logs; is supposed to work but did not - probably a version thing) to display the current logs saved in the mysql/var directory:
mysql> show master logs;
+---------------+
| Log_name |
+---------------+
| mysql-bin.002 |
| mysql-bin.003 |
| mysql-bin.004 |
+---------------+
3 rows in set (0.00 sec)
I then ran the purge master logs command to get rid of the deadwood:
mysql> purge master logs to 'mysql-bin.004';
Query OK, 0 rows affected (0.02 sec)
mysql> show master logs;
+---------------+
| Log_name |
+---------------+
| mysql-bin.004 |
+---------------+
1 row in set (0.00 sec)
It deleted the big files and we're out of the woods for disk space. I should probably set the max_binlog_size variable a little lower so it creates more, smaller logs so I don't reach a situation where I have a monstrous active log file and no old ones to purge.
Signing onto the slave server, I ran "show slave status;" at the mysql> prompt to show that the server was reading from the binary log called "mysql-bin.004" on the master.
Logging onto the master, I ran "show master logs;" at the mysql> prompt (show binary logs; is supposed to work but did not - probably a version thing) to display the current logs saved in the mysql/var directory:
mysql> show master logs;
+---------------+
| Log_name |
+---------------+
| mysql-bin.002 |
| mysql-bin.003 |
| mysql-bin.004 |
+---------------+
3 rows in set (0.00 sec)
I then ran the purge master logs command to get rid of the deadwood:
mysql> purge master logs to 'mysql-bin.004';
Query OK, 0 rows affected (0.02 sec)
mysql> show master logs;
+---------------+
| Log_name |
+---------------+
| mysql-bin.004 |
+---------------+
1 row in set (0.00 sec)
It deleted the big files and we're out of the woods for disk space. I should probably set the max_binlog_size variable a little lower so it creates more, smaller logs so I don't reach a situation where I have a monstrous active log file and no old ones to purge.
Wednesday, October 24, 2007
Perfect Storm Disk Replacement
I recently had a drive go bad in a Sun StorEdge 3510 FC JBOD array connected to a V490 running Solaris 10 with Solaris Volume Manager. The disk was part of a five-disk stripeset that was mirrored with another stripset.
It was *not* easy finding documentation for getting this done. Tools I'd used on other systems that had SCSI attached arrays and on systems with FC attached RAID arrays did not work. The combination of JBOD with FC on a 3510 managed with Solaris Volume Manager with an active hot spare made it interesting. So without further ado...
How To Replace a Failed Drive on a JBOD Sun StorEdge 3510 FC Array That Has Been Failed Over to a Hot Spare Managed by Volume Manager in Solaris 10 (whew!)
Here's the device with the bad disk c1t10d0s0 that was replaced with the hot spare from c1t11d0s0:
# metastat d15
d15: Mirror
Submirror 0: d16
State: Okay
Submirror 1: d17
State: Okay
Pass: 1
Read option: roundrobin (default)
Write option: parallel (default)
Size: 716634624 blocks (341 GB)
d16: Submirror of d15
State: Okay
Hot spare pool: hsp000
Size: 716634624 blocks (341 GB)
Stripe 0: (interlace: 256 blocks)
Device Start Block Dbase State Reloc Hot Spare
c1t4d0s0 20352 Yes Okay Yes
c1t3d0s0 20352 Yes Okay Yes
c1t2d0s0 20352 Yes Okay Yes
c1t1d0s0 20352 Yes Okay Yes
c1t0d0s0 20352 Yes Okay Yes
d17: Submirror of d15
State: Okay
Hot spare pool: hsp000
Size: 716634624 blocks (341 GB)
Stripe 0: (interlace: 256 blocks)
Device Start Block Dbase State Reloc Hot Spare
c1t9d0s0 20352 Yes Okay Yes
c1t8d0s0 20352 Yes Okay Yes
c1t7d0s0 20352 Yes Okay Yes
c1t6d0s0 20352 Yes Okay Yes
c1t10d0s0 20352 No Okay Yes c1t11d0s0
Device Relocation Information:
Device Reloc Device ID
c1t4d0 Yes id1,ssd@n20000011c6968cf9
c1t3d0 Yes id1,ssd@n20000011c6967f16
c1t2d0 Yes id1,ssd@n20000011c6968c7c
c1t1d0 Yes id1,ssd@n20000011c68baaed
c1t0d0 Yes id1,ssd@n20000011c6968ca1
c1t9d0 Yes id1,ssd@n20000011c6967e6e
c1t8d0 Yes id1,ssd@n20000011c68b0388
c1t7d0 Yes id1,ssd@n20000011c68deaaf
c1t6d0 Yes id1,ssd@n20000011c6969259
c1t11d0 Yes id1,ssd@n20000011c68bbb2d
I removed the meta database replicas that were on c1t10d0 but I'm not convinced I had to do that before continuing.
The cfgadm command can show the attachment point for the disk.
# cfgadm -al
Ap_Id Type Receptacle Occupant Condition
c0 scsi-bus connected configured unknown
c0::dsk/c0t0d0 CD-ROM connected configured unknown
c1 fc-private connected configured unknown
c1::22000011c68b0388 disk connected configured unknown
c1::22000011c68b5cb3 disk connected configured unknown
c1::22000011c68baaed disk connected configured unknown
c1::22000011c68bbb2d disk connected configured unknown
c1::22000011c68deaaf disk connected configured unknown
c1::22000011c6967e6e disk connected configured unknown
c1::22000011c6967f16 disk connected configured unknown
c1::22000011c6968c7c disk connected configured unknown
c1::22000011c6968ca1 disk connected configured unknown
c1::22000011c6968cf9 disk connected configured unknown
c1::22000011c6969259 disk connected configured unknown
c1::22000011c696a895 disk connected configured unknown
c1::225000c0ff086290 ESI connected configured unknown
c2 fc-private connected configured unknown
c2::500000e01127c191 disk connected configured unknown
c2::500000e01127c8a1 disk connected configured unknown
usb0/1 unknown empty unconfigured ok
usb0/2 unknown empty unconfigured ok
usb0/3 unknown empty unconfigured ok
usb0/4 unknown empty unconfigured ok
However, both the cfgadm and luxadm commands are unable to remove the drive since it's on a fiber loop and is a JBOD array.
# cfgadm -x replace_device c1::22000011c68b5cb3
cfgadm: Configuration operation not supported
# luxadm remove_device 22000011c68b5cb3
WARNING!!! Please ensure that no filesystems are mounted on these device(s).
All data on these devices should have been backed up.
Error: Invalid path. Device is not a SENA subsystem. - 22000011c68b5cb3.
Instead, use luxadm to offline the bad disk:
# luxadm -e offline /dev/rdsk/c1t10d0s2
Then devfsadm to remove the dev entries:
# devfsadm -Cv
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s0
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s1
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s2
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s3
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s4
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s5
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s6
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s7
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s0
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s1
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s2
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s3
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s4
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s5
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s6
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s7
The output from cfgadm now shows the device as unusable:
# cfgadm -al
Ap_Id Type Receptacle Occupant Condition
c0 scsi-bus connected configured unknown
c0::dsk/c0t0d0 CD-ROM connected configured unknown
c1 fc-private connected configured unknown
c1::22000011c68b0388 disk connected configured unknown
c1::22000011c68b5cb3 disk connected configured unusable
c1::22000011c68baaed disk connected configured unknown
c1::22000011c68bbb2d disk connected configured unknown
[...snip...]
Physically replace the device. In the 3510 JBOD array with the default boxid of zero (check the button hidden under the left plastic ear tab), the disk layout looks like this:
0 3 6 9
1 4 7 10
2 5 8 11
(0 to 11 counting down columns first then over rows)
When the disk is replaced, the devfsadm daemon should pick up the disk immediately and configure the dev entries. If not, try this to see what the problem is:
# luxadm -e port
/devices/pci@9,600000/SUNW,qlc@2/fp@0,0:devctl CONNECTED
/devices/pci@8,600000/SUNW,qlc@1/fp@0,0:devctl CONNECTED
Note: If you get a "NOT CONNECTED" error on the 3510 path, check cfgadm to see if the fiber connection is connected.
# cfgadm -al
Ap_Id Type Receptacle Occupant Condition
c0 scsi-bus connected configured unknown
c0::dsk/c0t0d0 CD-ROM connected configured unknown
c1 fc-private connected configured unknown
c1::22000011c68b0388 disk connected configured unknown
c1::22000011c68baaed disk connected configured unknown
c1::22000011c68bbb2d disk connected configured unknown
c1::22000011c68deaaf disk connected configured unknown
c1::22000011c6967e6e disk connected configured unknown
c1::22000011c6967f16 disk connected configured unknown
c1::22000011c6968c7c disk connected configured unknown
c1::22000011c6968ca1 disk connected configured unknown
c1::22000011c6968cf9 disk connected configured unknown
c1::22000011c6969259 disk connected configured unknown
c1::22000011c696a895 disk connected configured unknown
c1::225000c0ff086290 ESI connected configured unknown
c1::500000e014cb0282 disk connected configured unknown
c2 fc-private connected configured unknown
c2::500000e01127c191 disk connected configured unknown
c2::500000e01127c8a1 disk connected configured unknown
usb0/1 unknown empty unconfigured ok
usb0/2 unknown empty unconfigured ok
usb0/3 unknown empty unconfigured ok
usb0/4 unknown empty unconfigured ok
If the controller isn't there or is unconfigured try the following:
# cfgadm -c configure cx
If the drives appear with a condition set to "unusable" do the following using the pathname from the luxadm -e port command above:
# luxadm -e forcelip devices/pci@9,600000/SUNW,qlc@2/fp@0,0:devctl
Once the dev devices for the replaced drive are back in, use format to partition the new drive like the old one used to be. You can use the partition map from the hot spare as a template.
Once the drive is partitioned, add any database replicas that may have been on the original device (I should mention that I forgot to do that, so I'm not 100% sure that works), then do a metareplace to trigger the hot spare to go back to available and the replaced drive to start resyncing:
# metareplace -e d17 c1t10d0s0
Show progress with:
# metastat | grep %
Resync in progress: 73 % done
and see that the hot spare is available again with:
# metahs -i
# metahs -i
hsp000: 2 hot spares
Device Status Length Reloc
c1t11d0s0 Available 143349312 blocks Yes
c1t5d0s0 Available 143349312 blocks Yes
Device Relocation Information:
Device Reloc Device ID
c1t11d0 Yes id1,ssd@n20000011c68bbb2d
c1t5d0 Yes id1,ssd@n20000011c696a895
keywords: 3150 storedge storagetek solaris volume manager hot spare fc fiber channel jbod
It was *not* easy finding documentation for getting this done. Tools I'd used on other systems that had SCSI attached arrays and on systems with FC attached RAID arrays did not work. The combination of JBOD with FC on a 3510 managed with Solaris Volume Manager with an active hot spare made it interesting. So without further ado...
How To Replace a Failed Drive on a JBOD Sun StorEdge 3510 FC Array That Has Been Failed Over to a Hot Spare Managed by Volume Manager in Solaris 10 (whew!)
Here's the device with the bad disk c1t10d0s0 that was replaced with the hot spare from c1t11d0s0:
# metastat d15
d15: Mirror
Submirror 0: d16
State: Okay
Submirror 1: d17
State: Okay
Pass: 1
Read option: roundrobin (default)
Write option: parallel (default)
Size: 716634624 blocks (341 GB)
d16: Submirror of d15
State: Okay
Hot spare pool: hsp000
Size: 716634624 blocks (341 GB)
Stripe 0: (interlace: 256 blocks)
Device Start Block Dbase State Reloc Hot Spare
c1t4d0s0 20352 Yes Okay Yes
c1t3d0s0 20352 Yes Okay Yes
c1t2d0s0 20352 Yes Okay Yes
c1t1d0s0 20352 Yes Okay Yes
c1t0d0s0 20352 Yes Okay Yes
d17: Submirror of d15
State: Okay
Hot spare pool: hsp000
Size: 716634624 blocks (341 GB)
Stripe 0: (interlace: 256 blocks)
Device Start Block Dbase State Reloc Hot Spare
c1t9d0s0 20352 Yes Okay Yes
c1t8d0s0 20352 Yes Okay Yes
c1t7d0s0 20352 Yes Okay Yes
c1t6d0s0 20352 Yes Okay Yes
c1t10d0s0 20352 No Okay Yes c1t11d0s0
Device Relocation Information:
Device Reloc Device ID
c1t4d0 Yes id1,ssd@n20000011c6968cf9
c1t3d0 Yes id1,ssd@n20000011c6967f16
c1t2d0 Yes id1,ssd@n20000011c6968c7c
c1t1d0 Yes id1,ssd@n20000011c68baaed
c1t0d0 Yes id1,ssd@n20000011c6968ca1
c1t9d0 Yes id1,ssd@n20000011c6967e6e
c1t8d0 Yes id1,ssd@n20000011c68b0388
c1t7d0 Yes id1,ssd@n20000011c68deaaf
c1t6d0 Yes id1,ssd@n20000011c6969259
c1t11d0 Yes id1,ssd@n20000011c68bbb2d
I removed the meta database replicas that were on c1t10d0 but I'm not convinced I had to do that before continuing.
The cfgadm command can show the attachment point for the disk.
# cfgadm -al
Ap_Id Type Receptacle Occupant Condition
c0 scsi-bus connected configured unknown
c0::dsk/c0t0d0 CD-ROM connected configured unknown
c1 fc-private connected configured unknown
c1::22000011c68b0388 disk connected configured unknown
c1::22000011c68b5cb3 disk connected configured unknown
c1::22000011c68baaed disk connected configured unknown
c1::22000011c68bbb2d disk connected configured unknown
c1::22000011c68deaaf disk connected configured unknown
c1::22000011c6967e6e disk connected configured unknown
c1::22000011c6967f16 disk connected configured unknown
c1::22000011c6968c7c disk connected configured unknown
c1::22000011c6968ca1 disk connected configured unknown
c1::22000011c6968cf9 disk connected configured unknown
c1::22000011c6969259 disk connected configured unknown
c1::22000011c696a895 disk connected configured unknown
c1::225000c0ff086290 ESI connected configured unknown
c2 fc-private connected configured unknown
c2::500000e01127c191 disk connected configured unknown
c2::500000e01127c8a1 disk connected configured unknown
usb0/1 unknown empty unconfigured ok
usb0/2 unknown empty unconfigured ok
usb0/3 unknown empty unconfigured ok
usb0/4 unknown empty unconfigured ok
However, both the cfgadm and luxadm commands are unable to remove the drive since it's on a fiber loop and is a JBOD array.
# cfgadm -x replace_device c1::22000011c68b5cb3
cfgadm: Configuration operation not supported
# luxadm remove_device 22000011c68b5cb3
WARNING!!! Please ensure that no filesystems are mounted on these device(s).
All data on these devices should have been backed up.
Error: Invalid path. Device is not a SENA subsystem. - 22000011c68b5cb3.
Instead, use luxadm to offline the bad disk:
# luxadm -e offline /dev/rdsk/c1t10d0s2
Then devfsadm to remove the dev entries:
# devfsadm -Cv
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s0
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s1
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s2
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s3
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s4
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s5
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s6
devfsadm[3915]: verbose: removing file: /dev/dsk/c1t10d0s7
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s0
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s1
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s2
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s3
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s4
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s5
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s6
devfsadm[3915]: verbose: removing file: /dev/rdsk/c1t10d0s7
The output from cfgadm now shows the device as unusable:
# cfgadm -al
Ap_Id Type Receptacle Occupant Condition
c0 scsi-bus connected configured unknown
c0::dsk/c0t0d0 CD-ROM connected configured unknown
c1 fc-private connected configured unknown
c1::22000011c68b0388 disk connected configured unknown
c1::22000011c68b5cb3 disk connected configured unusable
c1::22000011c68baaed disk connected configured unknown
c1::22000011c68bbb2d disk connected configured unknown
[...snip...]
Physically replace the device. In the 3510 JBOD array with the default boxid of zero (check the button hidden under the left plastic ear tab), the disk layout looks like this:
0 3 6 9
1 4 7 10
2 5 8 11
(0 to 11 counting down columns first then over rows)
When the disk is replaced, the devfsadm daemon should pick up the disk immediately and configure the dev entries. If not, try this to see what the problem is:
# luxadm -e port
/devices/pci@9,600000/SUNW,qlc@2/fp@0,0:devctl CONNECTED
/devices/pci@8,600000/SUNW,qlc@1/fp@0,0:devctl CONNECTED
Note: If you get a "NOT CONNECTED" error on the 3510 path, check cfgadm to see if the fiber connection is connected.
# cfgadm -al
Ap_Id Type Receptacle Occupant Condition
c0 scsi-bus connected configured unknown
c0::dsk/c0t0d0 CD-ROM connected configured unknown
c1 fc-private connected configured unknown
c1::22000011c68b0388 disk connected configured unknown
c1::22000011c68baaed disk connected configured unknown
c1::22000011c68bbb2d disk connected configured unknown
c1::22000011c68deaaf disk connected configured unknown
c1::22000011c6967e6e disk connected configured unknown
c1::22000011c6967f16 disk connected configured unknown
c1::22000011c6968c7c disk connected configured unknown
c1::22000011c6968ca1 disk connected configured unknown
c1::22000011c6968cf9 disk connected configured unknown
c1::22000011c6969259 disk connected configured unknown
c1::22000011c696a895 disk connected configured unknown
c1::225000c0ff086290 ESI connected configured unknown
c1::500000e014cb0282 disk connected configured unknown
c2 fc-private connected configured unknown
c2::500000e01127c191 disk connected configured unknown
c2::500000e01127c8a1 disk connected configured unknown
usb0/1 unknown empty unconfigured ok
usb0/2 unknown empty unconfigured ok
usb0/3 unknown empty unconfigured ok
usb0/4 unknown empty unconfigured ok
If the controller isn't there or is unconfigured try the following:
# cfgadm -c configure cx
If the drives appear with a condition set to "unusable" do the following using the pathname from the luxadm -e port command above:
# luxadm -e forcelip devices/pci@9,600000/SUNW,qlc@2/fp@0,0:devctl
Once the dev devices for the replaced drive are back in, use format to partition the new drive like the old one used to be. You can use the partition map from the hot spare as a template.
Once the drive is partitioned, add any database replicas that may have been on the original device (I should mention that I forgot to do that, so I'm not 100% sure that works), then do a metareplace to trigger the hot spare to go back to available and the replaced drive to start resyncing:
# metareplace -e d17 c1t10d0s0
Show progress with:
# metastat | grep %
Resync in progress: 73 % done
and see that the hot spare is available again with:
# metahs -i
# metahs -i
hsp000: 2 hot spares
Device Status Length Reloc
c1t11d0s0 Available 143349312 blocks Yes
c1t5d0s0 Available 143349312 blocks Yes
Device Relocation Information:
Device Reloc Device ID
c1t11d0 Yes id1,ssd@n20000011c68bbb2d
c1t5d0 Yes id1,ssd@n20000011c696a895
keywords: 3150 storedge storagetek solaris volume manager hot spare fc fiber channel jbod
Subscribe to:
Posts (Atom)