Wednesday, May 7, 2014

Monitoring Linux server performance with procallator

Monitoring Linux server performance with procallator, 

I manage a fair number of Linux hosts, and like to keep tabs on how my systems are performing. One way I accomplish this is with procallator, which is a Perl script that collects performance data that can be graphed by orca. The graphs that orca produces are great awesome for trening server performance over time, and can be extremely valuable when debugging performance problems.

To setup procallator to collect performance data, you first need to retrieve the latest orca CVS snapshot from the orcaware snapshots directory (the procallator script is included with the orca snapshot, and the latest version contains a number of fixes). Once orca is downloaded, you will need to extract the tarball and run configure to modify the variables in the header of the procallator script:
$ tar xfj orca-snapshot-r529.tar.bz2
$ cd orca-snapshot-r529
$ ./configure –prefix=/opt/orca-r529 –with-html-dir=/opt/html
After the configure operation completes, you can install the procallator scripts with the Makefile’s install option:
$ make install
This will place the procallator perl script in $PREFIX/bin. To make sure the script starts at system boot, you can copy the $PREFIX/data_gathers/procallator/S99procallator script to /etc/rc3.d (or /etc/init.d depending on how you install your init scripts):
$ cp S99procallator /etc/rc3.d
Once these files are in place, you can start procallator by invoking the init script with the start option:
$ /etc/rc3.d/S99procallator start
This will start the procallator script as a daemon process, and the script will write performance data to the directory defined in the procallator script’s DEST_DIR variable every 5 minutes (this is tunable). The performance files will contain the name proccol-YYYY-MM-DD-INDEX, and one file will be produced each day. To graph the data in the procallator files, you can use orca and the procallator.cfg file that is in the $PREFIX/data_gathers/procallator directory. I placed a sample set of performance graphs on my website, and you can reference the article monitoring LDAP performance article for details on setting up orca to graph data. I digs me some procallator!

Why isn’t Oracle using huge pages on my Redhat Linux server?

Why isn’t Oracle using huge pages on my Redhat Linux server?

I am currently working on upgrading a number of Oracle RAC nodes from RHEL4 to RHEL5. After I upgraded the first node in the cluster, my DBA contacted me because the RHEL5 node was extremely sluggish. When I looked at top, I saw that a number of kswapd processes were consuming CPU:
$ top
top - 18:04:20 up 6 days,  3:22,  7 users,  load average: 14.25, 12.61, 14.41
Tasks: 536 total,   2 running, 533 sleeping,   0 stopped,   1 zombie
Cpu(s): 12.9%us, 19.2%sy,  0.0%ni, 20.9%id, 45.0%wa,  0.1%hi,  1.9%si,  0.0%st
Mem:  16373544k total, 16334112k used,    39432k free,     4916k buffers
Swap: 16777208k total,  2970156k used, 13807052k free,  5492216k cached

  PID USER      PR  NI  VIRT  RES  SHR S %CPU %MEM    TIME+  COMMAND                        
  491 root      10  -5     0    0    0 D 55.6  0.0  67:22.85 kswapd0                         
  492 root      10  -5     0    0    0 S 25.8  0.0  37:01.75 kswapd1                         
  494 root      11  -5     0    0    0 S 24.8  0.0  42:15.31 kswapd3                         
 8730 oracle    -2   0 8352m 3.5g 3.5g S  9.9 22.4 139:36.18 oracle                          
 8726 oracle    -2   0 8352m 3.5g 3.5g S  9.6 22.5 138:13.54 oracle                          
32643 oracle    15   0 8339m  97m  92m S  9.6  0.6   0:01.31 oracle                          
  493 root      11  -5     0    0    0 S  9.3  0.0  43:11.31 kswapd2                         
 8714 oracle    -2   0 8352m 3.5g 3.5g S  9.3 22.4 137:14.96 oracle                          
 8718 oracle    -2   0 8352m 3.5g 3.5g S  8.9 22.3 137:01.91 oracle                          
19398 oracle    15   0 8340m 547m 545m R  7.9  3.4   0:05.26 oracle                          
 8722 oracle    -2   0 8352m 3.5g 3.5g S  7.6 22.5 139:18.33 oracle              
The kswapd process is responsible for scanning memory to locate free pages, and scheduling dirty pages to be written to disk. Periodic kswapd invocations are fine, but seeing kswapd continuosly appearing in the top output is a really really bad thing. Since this host should have had plenty of free memory, I was perplexed by the following output (the free output didn’t match up with the values on the other nodes):
$ free
             total       used       free     shared    buffers     cached
Mem:      16373544   16268540     105004          0       1520    5465680
-/+ buffers/cache:   10801340    5572204
Swap:     16777208    2948684   13828524
To start debugging the issue, I first looked at ipcs to see how much shared memory the database allocated. In the output below, we can see that there is a 128MB and a 8GB shared memory segment allocated:
$ ipcs -a
------ Shared Memory Segments --------
key        shmid      owner      perms      bytes      nattch     status      
0x62e08f78 0          oracle    640        132120576  16                      
0xdd188948 32769      oracle    660        8592031744 87         
The first segment is dedicated to the Oracle ASM instance, and the second to the actual database. When I checked the number of huge pages allocated to the machine, I saw something a bit odd:
$ grep Huge cat /proc/meminfo
HugePages_Total:  4106
HugePages_Free:   4051
HugePages_Rsvd:      8
Hugepagesize:     2048 kB
While our DBA had set vm.nr_hugepages to a sufficiently large value in /etc/syscl.conf, the database was utilizing a very small portion of them. This meant that the database was being allocated out of non huge page memory (Linux dedicates memory to the huge page area, and it is wasted if nothing utilizes it), and inactive pages were being paged out to disk since the database wasn’t utilizing the huge page area we reserved for it . After a bit of bc’ing (I love doing my calculations with bc), I noticed that the total amount of memory allocated to huge pages was 8610906112 bytes:
$ grep vm.nr_hugepages /etc/sysctl.conf
vm.nr_hugepages=4106
$ bc
4106*(1024*1024*2)
8610906112
If we add the totals from the two shared memory segments above:
$ bc
8592031744+132120576
8724152320
We can see that we don’t have enough huge page memory to support both shared memory segments. Yikes! After adjusting vm.nr_hugepages to account for both databases, the system no longer swapped and database performance increased. This debugging adventure taught me a couple of things:
1. Double check system values people send you
2. Solaris does a MUCH better job of handling large page sizes (huge pages are used transparently)
3. The Linux tools for investigating huge page allocations are severely lacking
4. Oracle is able to allocate a continuos 8GB chunk of shared memory on RHEL5, but not RHEL4 (I need to do some research to find out why)
Hopefully more work will go into the Linux huge page implementation, and allow me to scratch the second and third items off of my list. Viva la problem resolution!

Getting an accurate view of process memory usage on Linux hosts

Getting an accurate view of process memory usage on Linux hosts

Having debugged a number of memory-related issues on Linux, one thing I’ve always wanted was a tool to display proportional memory usage. Specifically, I wanted to be able to see how much memory was unique to a process, and have an equal portion of shared memory (libraries, SMS, etc.) added to this value. My wish came true a while back when I discovered the smem utility. When run without any arguments, smem will give you the resident set size (RSS), the unique set size (USS) and the proportional set size (PSS) which is the unique set size plus a portion of the shared memory that is being used by this process. This results in output similar to the following:
$ smem -r
  PID User     Command                         Swap      USS      PSS      RSS 
 3636 root     /usr/lib/virtualbox/Virtual        0  1151596  1153670  1165568 
 3678 matty    /usr/lib64/firefox-3.5.9/fi        0   189628   191483   203028 
 5779 root     /usr/bin/python /usr/bin/sm        0    38584    39114    40368 
 1847 root     /usr/bin/Xorg :0 -nr -verbo        0    34024    35874    92504 
 4103 matty    pidgin                             0    19364    21072    32412 
 3825 matty    gnome-terminal                     0    12388    13242    21992 
 3404 matty    python /usr/share/system-co        0    11596    12622    19216 
 3710 matty    gnome-screensaver                  0     9872    10287    14640 
 3283 matty    nautilus                           0     7104     8373    18484 
 3263 matty    gnome-panel                        0     5828     6731    15780 
To calculate the portion of shared memory that is being used by each process, you can add up the shared memory per process (you would probably index this by the type of shared resource), the number of processes using these pages, and then divide the two values to get a proportional value of shared memory per process. This is a very cool utility, and one that gets installed on all of my systems now!

Installing QLogic drivers on CentOS Linux hosts

Installing QLogic drivers on CentOS Linux hosts

I have a couple of hosts in my lab with QLA2342 HBAs, and use the drivers from QLogic instead of the drivers that come with the kernel. There are a number of reasons for this, but I’ll save that explanation for a future post. To install the QLogic drivers on a CentOS 5.3 host, you will first need to download the qlafc package from QLogic’s website. Once you retrieve the bits, you can run the qlinstall utility to build and install a qlaXxxx driver that works with the kernel you are running:
$ tar xfvz qlafc-linux-8.02.14.01-1-install.tgz
$ cd qlafc-linux-8.02.14.01-1-install
$ ls -la
total 11772
drwxr-xr-x 4 root  root     4096 Jul 10  2008 .
drwx------ 3 matty matty    4096 Apr 30 07:45 ..
drwxr-xr-x 4 root  root     4096 Jul 10  2008 agents
drwxr-xr-x 7 root  root     4096 Apr 30 07:47 LinuxTools
-rw-r--r-- 1 root  root  2376930 Jul 10  2008 qla2xxx-v8.02.14.01-1.noarch.rpm
-rwxr-xr-x 1 root  root   146155 Jul 10  2008 qlinstall
-rw-r--r-- 1 root  root     7321 Jul 10  2008 ql-pci.ids
-rw-r--r-- 1 root  root    14119 Jul 10  2008 README.qlinstall.txt
-rw-r--r-- 1 root  root  2905583 Jul 10  2008 scli-1.7.1-23.i386.rpm
-rw-r--r-- 1 root  root  3736548 Jul 10  2008 scli-1.7.1-23.ia64.rpm
-rw-r--r-- 1 root  root  2792038 Jul 10  2008 scli-1.7.1-23.ppc64.rpm
-rwxr-xr-x 1 root  root    18465 Jul 10  2008 set_driver_param


$ ./qlinstall
#*********************************************************#
#           SANsurfer Driver Installer for Linux          #
#             Installer Version:  1.01.00pre21        #
#*********************************************************#

Kernel version: 2.6.18-128.1.6.el5
Distribution: CentOS release 5.3 (Final)

Found following QLogic Adapter in the system
    1. QLA2342
Installation will begin for following driver
    1. qla2xxx version: v8.02.14.01


Unloading any loaded drivers
Unloaded module qla2xxx

Installing Driver...
Preparing...                ##################################################
qla2xxx                     ##################################################

QLA2XXX -- Building the qla2xxx driver, please wait...
Installing intermodule.ko in /lib/modules/2.6.18-128.1.6.el5/kernel/kernel/
 


QLA2XXX -- Installing the qla2xxx modules to 
/lib/modules/2.6.18-128.1.6.el5/kernel/drivers/scsi/qla2xxx/...

Setting up QLogic HBA API library...
Please make sure the /usr/lib/libqlsdm.so file is not in use.
Installing 32bit api binary for x86_64.
Installing 64bit api binary for x86_64.
Library 4.00 build12 already installed at /usr/lib/libqlsdm.so.
Done.

Loading module qla2xxx_conf version: v8.02.14.01....
Loaded module qla2xxx_conf
Loading module qla2xxx version: v8.02.14.01....
Loaded module qla2xxx


Building default persistent binding using SCLI
 
Saved copy of /etc/modprobe.conf as
/usr/src/qlogic/v8.02.14.01-1/backup/modprobe.conf-2.6.18-128.1.6.el5-043009-082949.bak

Saved copy of /boot/initrd-2.6.18-128.1.6.el5.img as
/boot/initrd-2.6.18-128.1.6.el5_QLI.org
QLA2XXX -- Rebuilding ramdisk image...
Ramdisk created.

Reloading the QLogic FC HBA drivers....
Unloaded module qla2xxx
Loading module qla2xxx_conf version: v8.02.14.01....
Loaded module qla2xxx_conf
Loading module qla2xxx version: v8.02.14.01....
Loaded module qla2xxx


Target Information on all HBAs:
==============================
 --------------------------------------------------------------------------------
HBA Instance 0: QLA2342 Port 1 WWPN 21-00-00-1B-32-04-86-C3 PortID 01-12-00
--------------------------------------------------------------------------------
No device connected to selected HBA (Instance 0)!
--------------------------------------------------------------------------------
HBA Instance 1: QLA2342 Port 2 WWPN 21-01-00-1B-32-24-86-C3 PortID 01-13-00
--------------------------------------------------------------------------------
No device connected to selected HBA (Instance 1)!
Installing the qlinstall-autoload script in /etc/init.d/ 

#*********************************************************#
#               INSTALLATION SUCCESSFUL!!                 #
#    SANsurfer Driver installation for Linux completed    #
#*********************************************************#


Once the driver is installed and loaded, you can verify the version by running the modinfo utility:
$ modinfo qla2xxx | head -10
filename:       /lib/modules/2.6.18-128.1.6.el5/kernel/drivers/scsi/qla2xxx/qla2xxx.ko
version:        8.02.14.01
license:        GPL
description:    QLogic Fibre Channel HBA Driver
author:         QLogic Corporation
srcversion:     90F727BC62E8CB20433AB36
alias:          pci:v00001077d00002532sv*sd*bc*sc*i*
alias:          pci:v00001077d00005432sv*sd*bc*sc*i*
alias:          pci:v00001077d00005422sv*sd*bc*sc*i*
alias:          pci:v00001077d00008432sv*sd*bc*sc*i*


QLogic includes a number of nifty utilities with their driver bundle, all of which I plan to blog about in future posts.

QLogic ISP2532-based 8Gb Fibre Channel


QLogic ISP2532-based 8Gb Fibre Channel


You shouldn't need to do anything as the driver is already available as part of the CentOS supplied kernel and modules. To check this, I looked up ISP2532 in /usr/share/hwdata/pci.ids and found that it has PCI Ventodr:Device ID 1077:2532. Then I ran

CODE: SELECT ALL
$ grep 1077 /lib/modules/2.6.18-348.1.1.el5xen/modules.* | grep 2532
/lib/modules/2.6.18-348.1.1.el5xen/modules.alias:alias pci:v00001077d00002532sv*sd*bc*sc*i* qla2xxx


This shows that the device is supported by the qla2xxx module. If running `lsmod | grep qla2xxx` on the running system returns no output, then try running `modprobe qla2xxx` and then look in /var/log/messages to see what appears at the end of the file. You might want to check the output from `lspci -nn` to make sure that your adapter has the same PCI Vendor:Device ID that I looked for.

Linux Bass-files Fiber Channel Multipath configuration

Linux Bass-files Fiber Channel Multipath configuration

INSTALL Red Hat 5.3                     Update March 31, 2009
-------------------

Make sure the fibre hba cables are UNPLUGGED. The Red Hat installer gets
confused with the multiple paths to the 6140 StorageTek disk array.
There are 4 paths to each disk drive LUN.  We think this is because
the existing LUNS that are presented to the server are lvm logical
volumes. The server see's multiple paths to the same lvm physical
volume and this causes problems. Once the system sees the lvm physical
disks as one device things work properly. That is the qlogic and rdac
drivers need to be configured before the hba fibre cables are plugged
in. The RDAC drivers make the multiple paths to the LUNS's look like
one device. Note you can also disable the port in the Brocade switch if
you like but that is much more complicated. You can also use the Brocade
java gui to disable the ports.

Do a normal Red Hat 5 install. Do not plug the hba cables in yet! Do not
use the Red Hat multipath software. By default the /etc/multipath.conf
file has all disk devices disabled for use as multipath devices.

Install the Qlogic drivers first then Sun RDAC/MPP drivers.

NOTE: The Red Hat supplied qlogic drivers do not work,
      you need to use the Qlogic supplied drivers!

Get the latest Qlogic drivers at:

http://support.qlogic.com/

Click the downloads tab at the top of the page. Click on Sun under the
OEM MODELS heading.

Look for model SG-XPCIE2FC-QF4, click on "Software for:" Linux URL just
below the table!

Download the latest "Linux FC driver for 2.6 kernel (x86/x64/IA64)"
tarball file, for example:

        qla2xxx-v8.02.14_01-dist.tgz


There is also a "SANsurfer Linux Driver Installer (x86/x64/IA64)" download
that contains the SANsurfer command line utility and the driver. Do not
use this method for installing. Download the "SANsurfer
CLI (x86/x64)" rpm by itself and install:

        scli-1.7.1-23.i386.rpm.gz

gunzip scli-1.7.1-23.i386.rpm.gz
rpm -vih scli-1.7.1-23.i386.rpm

The scli command gets install in:

        /opt/QLogic_Corporation/SANsurferCLI/scli

Use the scli command in this directory to probe information about the
qlogic card WWPN World Wide Port Nodes. Use this utility after
you install the Qlogic driver.

Install the Qlogic provided driver. We will be using the Linux FC
driver for 2.6 download which is a compressed tar file. Do not use the
"qlinstall" command that comes with the "SANsurfer Linux Driver Installer
(x86/x64/IA64)" tarball! Here are the steps to instll the Qlogic driver
using the version as of this writing:

copy the downloaded driver software to /var/tmp and unpack:

# tar zxvf qla2xxx-v8.02.14_01-dist.tgz

The Qlogic installation instructions are in the file:

        README.qla2xxx


# cd qlogic
# ./drvrsetup
# cd qla2xxx-8.02.14
# ./extras/build.sh install

--- output example ---

bash-3.2# ./extras/build.sh install

QLA2XXX -- Building the qla2xxx driver, please wait...
Installing intermodule.ko in /lib/modules/2.6.18-92.1.10.el5/kernel/kernel/
QLA2XXX -- Build done.

QLA2XXX -- Installing the qla2xxx modules to 
/lib/modules/2.6.18-92.1.10.el5/kernel/drivers/scsi/qla2xxx/...

--- end output example ---

Place the following option in /etc/modprob.conf The RDAC driver cannot
co-exist with an HBA-level failover driver. Use echo or vi to update
/etc/modprobe.conf:

# echo "options qla2xxx ql2xfailover=0"  >> /etc/modprobe.conf

Optionaly make a new boot/intrd-*.img ram disk image.  You only need
to do this if you are troubleshooting before you install the Sun RDAC
drivers! The Sun RDAC build makes a new initrd-*.img file also.

# cd /boot

Save old image:

# cp initrd-`uname -r`.img initrd-`uname -r`.img.bak

Make new image to load qlogic driver on boot:

# mkinitrd -f initrd-`uname -r`.img `uname -r`

Reboot system, do not plug fibre cable in yet! Skip to the next step,
installing the RDAC drivers and do not reboot if you do not need to
troubleshoot the LUNS.

# init 6

Use the Qlogic scli command above to troubleshoot the
LUNS/StorageTek 6140.


INSTALL SUN RDAC/MPP
--------------------
Now get and install the Sun RDAC/MPP driver, Redundant Disk Array
Controller Driver (a.k.a. Multi-Path Proxy Driver or MPP)

AS of this writing you need to contact sun to get the latest linux driver.
Below is the output of the make install command.

# tar zxvf rdac-LINUX-09.02.C2.13-source.tar.gz
# cd linuxrdac-09.02.C2.13
# make
# make uninstall
# make install

If you do not do a "make uninstall" you will get this error if upgrading
to the kernel, the builds needs to remove all old files:

The system has old MPP driver package installed.
Please do "make uninstall" in the current directory before installing the new one.
make: *** [hbacheck] Error 1

Do as it says "make uninstall" then do "make install".
Also answer yes to the following question:

Host Adapters from different supported vendors co-exists on your system.
Please make sure that only one supported model of HBA is connected to Storage Array.
Do you want to continue (yes or no) ? yes

The installer will create a new mpp-kernel_release.img ram disk. Update
grub.conf to use this ramd disk, for example:

# vi /etc/grub.conf or /boot/grub/grub.conf

default=0
timeout=5
serial --unit=0 --speed=9600
terminal --timeout=5 serial console
title Red Hat Enterprise Linux Server SUN-RDAC (2.6.18-128.1.1.el5)
        root (hd0,0)
        kernel /vmlinuz-2.6.18-128.1.1.el5 ro root=/dev/sdasysvg/root
        console=ttyS0,9600n1 rhgb quiet
        initrd /mpp-2.6.18-128.1.1.el5.img

You can save the old kernel entries, make this entry the default boot
entry. This way you can boot the non RDAC kernel if you have problems.

Halt the system, "init 0". You can now attach the hba fibre cables.
Bring up the system. The lsmod command should show the qla drivers get
loaded and should look like this:

# lsmod|grep ql
qla2xxx              1009580  1 
qla2xxx_conf          335368  1 
intermodule            37508  2 qla2xxx,qla2xxx_conf
scsi_mod              188665  10 qla2xxx,sr_mod,usb_storage,mptsas,mptscsih,scsi_transport_sas,,sg,sd_mod


The disk LUNS will get presented as a single /dev/sd? device. The RDAC
driver coalesces the multiple paths to the disk array as one device.
You can use "fdisk -l" to list your disk devices that were created from
the LUNS. You can also use the RDAC command /opt/mpp/lsvdev to see
the Lun ot disk  mappings, very handy.

If you were using lvm logical volumes on the LUNS run the following
commands so the os will see the lvm physical volumes:

Scan for lvm physical volumes
# /sbin/pvscan

Verify the list of lvm physical volumes:
# /usr/sbin/pvs
# /usr/sbin/lvs -o +devices

You probably will have to turn the "available" attribute on for the
logical volumes after a fresh install. You can see if the lvm logic
device exists in /dev, for example "ls /dev/volume_group_name/*", If he
volume group device does not exist enable the available attribute for
all the volume groups with:

# lvm vgchange -ay volume_group volume_group1 ...

This should create the logical volume groups and volume device:

        /dev/volume_groupx/logical_volume_name

Again, use "ls /dev/volume_group_name/*" to see if the volume groups
got created

You can then mount the logical volume:
        mount /dev/volume_groupx/logical_volume_name /mount_point

Update the /etc/fstab for all your mounts and mount them.

Reboot the system and verify the boot process and mounts.


UPGRADING TO A NEW KERNEL
-------------------------

You can leave the fibre cables plugged in at this point. The Qlogic/RDAC
drivers only install to the kernel that is currently running.

Run yum to update your system and its kernel.

Comment out all StorageTek 6140 file systems that are mounted!  If you
forget to do this and boot the new kernel fsck will complain and drop
you into single user mode first asking you to enter the root password. If
you accidentally do this here is how to fix:

Enter root password and remount / as a read/write file system with:

        /bin/mount -o remount,rw /

edit /etc/fstab with vi and comment out the StorageTek file systems,
or execut "touch /fastboot" which creates and empty file that fsck
checks for and will bypass the fsck check.

Halt the system and remove the fibre cables from the system. Or user
the Brocade guid to disable the ports for the host.

Boot to the new kernel and install the Qlogic and RDAC drivers as
described above.

After checking the LUNS and volumes are present as described above
uncomment the /etc/fstab SotragTek mounts and re-mount the file systems.


QLOGIC SANSURFER UTILITY
------------------------

Get the qlafc-linux-8.02.08-1-install.tgz as described above and
unpack it:

tar ztvf qlafc-linux-8.02.08-1-install.tgz
cd qlafc-linux-8.02.08-1-install


Install the rpm for your platform, for example:

rpm -ivh scli-1.7.1-18.i386.rpm

The utility gets install at:

/opt/QLogic_Corporation/SANsurferCLI/scli



Some helful commands and file info:
----------------------------------

man RDAC, explains RDAC driver 

/opt/mpp/lsvdev command that displays lun to sd device mappings.

/proc/scsi/mpp/3, (or maybe some other number),  file that contains lun
and controller info.

mppUtil -S command, this will print all the multipaths, read mppUtil man page.

mppBusRescan command  run after mapping new lun so host registers it
on a live sytem if you do not want to reboot.

vgdisplay: Volume group not activated

vgdisplay: Volume group not activated


Here is the full message being shown to me by the OS after invoking the vgdisplay command:

vgdisplay: Volume group not activated.
vgdisplay: Cannot display volume group "/dev/vg10".

If the server is NOT part of a cluster, use

# vgchange -a y vg10

# strings /etc/lvmtab

# more /etc/fstab

# mount <info from fstab>



  

  DAILY CHECKLIST Oracle Database instance is running or not select name,open_mode from V$database; Note: Check the Oracle databases are ru...