Showing posts with label RAC. Show all posts
Showing posts with label RAC. Show all posts

Thursday, April 23, 2015

Oracle Cluster Health Monitor (CHM) using large amount of space (crfclust.bdb)

Last night my rac 2 node server went down for OS patcing and rebooted but all CRS resources not coming up on both the node after node reboots:

conn as root user and check all resources

[root@oradev11 bin]# ./crsctl stat res -t -init
--------------------------------------------------------------------------------
NAME           TARGET  STATE        SERVER                   STATE_DETAILS
--------------------------------------------------------------------------------
Cluster Resources
--------------------------------------------------------------------------------
ora.asm
      1        ONLINE  OFFLINE
ora.cluster_interconnect.haip
      1        ONLINE  OFFLINE
ora.crf
      1        ONLINE  OFFLINE
ora.crsd
      1        ONLINE  OFFLINE
ora.cssd
      1        ONLINE  OFFLINE
ora.cssdmonitor
      1        ONLINE  ONLINE       oradev11
ora.ctssd
      1        ONLINE  OFFLINE
ora.diskmon
      1        ONLINE  OFFLINE
ora.drivers.acfs
      1        ONLINE  ONLINE       oradev11
ora.evmd
      1        ONLINE  OFFLINE
ora.gipcd
      1        ONLINE  OFFLINE
ora.gpnpd
      1        ONLINE  OFFLINE
ora.mdnsd
      1        ONLINE  OFFLINE                               STARTING


CRS alert log says:

[root@oradev11 ] # cd $GRID_HOME/log/hostname
[root@oradev11 oradev11]# tail -50f alertoradev11.log

o/p trimmed………

2015-04-22 20:06:01.173:
[/u01/app/11.2.0.4/grid/bin/oraagent.bin(24990)]CRS-5818:Aborted
2015-04-22 20:06:05.177:
[ohasd(12696)]CRS-2757:Command 'Start' timed out waiting for response from the resource 'ora.mdnsd'. Details at (:CRSPE00111:) {0:0:2} in /u01/app/11.2.0.4/grid/log/sl73vmhasd/ohasd.log.
2015-04-22 20:06:05.658:
[/u01/app/11.2.0.4/grid/bin/oraagent.bin(25614)]CRS-0037:An error occurred while attempting to write to file "/u01/app/11.2.0.4/grid/log/oradev11/agent/ohasd/oraagenagent_grid.log". Additional diagnostics: LFI-00004: Call to lfibwrt() failed.
LFI-01518: write() failed(OSD return value = 28) in slfiwl.

2015-04-22 20:06:05.659:
[/u01/app/11.2.0.4/grid/bin/oraagent.bin(25614)]CRS-0004:logging terminated for the process. log file: "/u01/app/11.2.0.4/grid/log/oradev11/agent/ohasd/oraagent_gridgrid.log"
2015-04-22 20:06:06.176:
[/u01/app/11.2.0.4/grid/bin/oraagent.bin(25631)]CRS-0037:An error occurred while attempting to write to file "/u01/app/11.2.0.4/grid/log/oradev11/agent/ohasd/oraagenagent_grid.log". Additional diagnostics: LFI-00004: Call to lfibwrt() failed.
LFI-01518: write() failed(OSD return value = 28) in slfiwl.

2015-04-22 20:06:06.176:
[/u01/app/11.2.0.4/grid/bin/oraagent.bin(25631)]CRS-0004:logging terminated for the process. log file: "/u01/app/11.2.0.4/grid/log/oradev11/agent/ohasd/oraagent_gridgrid.log"
2015-04-22 20:06:06.272:
[gpnpd(25644)]CRS-0037:An error occurred while attempting to write to file "/u01/app/11.2.0.4/grid/log/oradev11/gpnpd/gpnpd.log". Additional diagnostics: LFI-00004: ibwrt() failed.
LFI-01518: write() failed(OSD return value = 28) in slfiwl.

2015-04-22 20:06:06.272:
[gpnpd(25644)]CRS-0004:logging terminated for the process. log file: "/u01/app/11.2.0.4/grid/log/oradev11/gpnpd/gpnpd.log"
2015-04-22 20:06:09.314:
[gpnpd(25644)]CRS-2329:GPNPD on node oradev11 shutdown.
2015-04-22 20:08:06.226:
[/u01/app/11.2.0.4/grid/bin/oraagent.bin(25631)]CRS-5818:Aborted command 'start' for resource 'ora.gpnpd'. Details at (:CRSAGF00113:) {0:0:2} in /u01/app/11.2.0.4/grid/logbd001/agent/ohasd/oraagent_grid/oraagent_grid.log.
2015-04-22 20:08:10.229:
[ohasd(12696)]CRS-2757:Command 'Start' timed out waiting for response from the resource 'ora.gpnpd'. Details at (:CRSPE00111:) {0:0:2} in /u01/app/11.2.0.4/grid/log/sl73vmhasd/ohasd.log.
2015-04-22 20:08:10.710:
[/u01/app/11.2.0.4/grid/bin/oraagent.bin(26582)]CRS-0037:An error occurred while attempting to write to file "/u01/app/11.2.0.4/grid/log/oradev11/agent/ohasd/oraagenagent_grid.log". Additional diagnostics: LFI-00004: Call to lfibwrt() failed.
LFI-01518: write() failed(OSD return value = 28) in slfiwl.

2015-04-22 20:08:10.710:
[/u01/app/11.2.0.4/grid/bin/oraagent.bin(26582)]CRS-0004:logging terminated for the process. log file: "/u01/app/11.2.0.4/grid/log/oradev11/agent/ohasd/oraagent_gridgrid.log"
2015-04-22 20:08:11.280:
[/u01/app/11.2.0.4/grid/bin/oraagent.bin(26604)]CRS-0037:An error occurred while attempting to write to file "/u01/app/11.2.0.4/grid/log/oradev11/agent/ohasd/oraagenagent_grid.log". Additional diagnostics: LFI-00004: Call to lfibwrt() failed.
LFI-01518: write() failed(OSD return value = 28) in slfiwl.

2015-04-22 20:08:11.280:
[/u01/app/11.2.0.4/grid/bin/oraagent.bin(26604)]CRS-0004:logging terminated for the process. log file: "/u01/app/11.2.0.4/grid/log/oradev11/agent/ohasd/oraagent_gridgrid.log"
2015-04-22 20:08:11.347:
[mdnsd(26617)]CRS-0037:An error occurred while attempting to write to file "/u01/app/11.2.0.4/grid/log/oradev11/mdnsd/mdnsd.log". Additional diagnostics: LFI-00004: ibwrt() failed.
LFI-01518: write() failed(OSD return value = 28) in slfiwl.

2015-04-22 20:08:11.347:
[mdnsd(26617)]CRS-0004:logging terminated for the process. log file: "/u01/app/11.2.0.4/grid/log/oradev11/mdnsd/mdnsd.log"
2015-04-22 20:08:11.351:
[mdnsd(26617)]CRS-5602:mDNS service stopping by request.

After so much of time spending on troubleshooting I checked the space on server  and then released it is because of space issue on a mount point where my GRID home located, due to which CRS resources are not coming up

[root@oradev11 bin]# df -h
Filesystem            Size  Used Avail Use% Mounted on
/dev/mapper/rootvg-rootlv
                      5.8G  1.8G  3.8G  33% /
tmpfs                 3.0G     0  3.0G   0% /dev/shm
/dev/sda1             190M   86M   95M  48% /boot
/dev/mapper/rootvg-homelv
                      2.0G  9.2M  1.8G   1% /home
/dev/mapper/rootvg-optlv
                      9.8G  2.0G  7.3G  22% /opt
/dev/mapper/rootvg-securlv
                      1.5G  211M  1.2G  16% /opt/security
/dev/mapper/rootvg-tmplv
                      2.0G  375M  1.5G  21% /tmp
/dev/mapper/rootvg-varlv
                      9.8G  1.1G  8.2G  12% /var
/dev/mapper/datavg-gridbaselv
                       50G   49G     0 100% /u01/app
/dev/mapper/datavg-rdbmsbaselv
                       50G  4.8G   42G  11% /u01/app/oracle
/dev/mapper/datavg-adrrepolv
                       50G  2.6G   45G   6% /oratrace
/dev/mapper/datavg-oemagentlv
                       20G  651M   18G   4% /u01/app/emagent
/dev/mapper/datavg-gglv
                       50G   52M   47G   1% /gg
/dev/mapper/datavg-dbawslv
                       99G   16G   79G  17% /oraworkspace
/dev/mapper/datavg-auditfslv
                       50G  230M   47G   1% /oradbaudit
/dev/mapper/datavg-dbtoolslv
                      9.8G   86M  9.2G   1% /oratools

Checking to see if i can delete anything on /u01/app mount point and i see "crfclust.bdb" is consuming much space then any other

[root@oradev11 bin]# cd ../crf/db
[root@oradev11 db]# ls -lrht
total 4.0K
drwxr-x--- 2 root oinstall 4.0K Apr 22 20:45 oradev11
[root@oradev11 db]# cd oradev11

[root@oradev11 oradev11]# ls -lrth
total 38G
-rw-r--r-- 1 root root 1.1M Sep  8  2014 08-SEP-2014-09:24:06.txt
-rw-r--r-- 1 root root 1.9M Sep  8  2014 08-SEP-2014-10:07:28.txt
-rw-r--r-- 1 root root 1.2M Sep  8  2014 08-SEP-2014-10:20:00.txt
-rw-r----- 1 root root 8.0K Nov 20 09:44 repdhosts.bdb
-rw-r--r-- 1 root root  74K Mar  9 10:53 09-MAR-2015-10:53:37.txt
-rw-r--r-- 1 root root 856K Mar  9 10:56 09-MAR-2015-10:56:42.txt
-rw-r--r-- 1 root root  77K Mar 13 19:21 13-MAR-2015-19:21:26.txt
-rw-r--r-- 1 root root 218K Mar 13 19:21 13-MAR-2015-19:21:44.txt
-rw-r----- 1 root root  16M Apr 22 12:19 log.0000007983
-rw-r----- 1 root root  24K Apr 22 20:42 __db.001
-rw-r--r-- 1 root root 115M Apr 22 20:42 oradev11.ldb
-rw-r----- 1 root root 8.0K Apr 22 20:43 crfconn.bdb
-rw-r--r-- 1 root root 777K Apr 22 20:45 22-APR-2015-20:45:53.txt
-rw-r----- 1 root root  56K Apr 22 20:56 __db.006
-rw-r----- 1 root root 392K Apr 22 20:56 __db.002
-rw-r----- 1 root root 812M Apr 22 20:56 crfloclts.bdb
-rw-r----- 1 root root 668M Apr 22 20:56 crfcpu.bdb
-rw-r----- 1 root root 743M Apr 22 20:56 crfalert.bdb
-rw-r----- 1 root root 526M Apr 22 20:56 crfts.bdb
-rw-r----- 1 root root 607M Apr 22 20:56 crfhosts.bdb
-rw-r----- 1 root root  34G Apr 22 20:56 crfclust.bdb
-rw-r----- 1 root root  16M Apr 22 20:56 log.0000007984
-rw-r----- 1 root root 1.2M Apr 22 20:56 __db.005
-rw-r----- 1 root root 2.1M Apr 22 20:56 __db.004
-rw-r----- 1 root root 2.6M Apr 22 20:56 __db.003

From the above output I see only “crfclust.bdb” is consuming lot of space, then I followed the steps given in the oracle doc to free up the space on the server


Stop ora.crf ……….

[root@oradev11 bin]# ./crsctl stop res ora.crf -init
CRS-2673: Attempting to stop 'ora.crf' on 'oradev11'
CRS-2677: Stop of 'ora.crf' on 'oradev11' succeeded

[root@oradev11 oradev11]# rm crfclust.bdb

[root@oradev11 oradev11]# df -h
Filesystem            Size  Used Avail Use% Mounted on
/dev/mapper/rootvg-rootlv
                      5.8G  1.8G  3.8G  33% /
tmpfs                 3.0G  854M  2.2G  28% /dev/shm
/dev/sda1             190M   86M   95M  48% /boot
/dev/mapper/rootvg-homelv
                      2.0G  9.2M  1.8G   1% /home
/dev/mapper/rootvg-optlv
                      9.8G  2.0G  7.3G  22% /opt
/dev/mapper/rootvg-securlv
                      1.5G  211M  1.2G  16% /opt/security
/dev/mapper/rootvg-tmplv
                      2.0G  376M  1.5G  21% /tmp
/dev/mapper/rootvg-varlv
                      9.8G  1.1G  8.2G  12% /var
/dev/mapper/datavg-gridbaselv
                       50G   13G   34G  28% /u01/app
/dev/mapper/datavg-rdbmsbaselv
                       50G  4.8G   42G  11% /u01/app/oracle
/dev/mapper/datavg-adrrepolv
                       50G  2.6G   45G   6% /oratrace
/dev/mapper/datavg-oemagentlv
                       20G  651M   18G   4% /u01/app/emagent
/dev/mapper/datavg-gglv
                       50G   52M   47G   1% /gg
/dev/mapper/datavg-dbawslv
                       99G   16G   79G  17% /oraworkspace
/dev/mapper/datavg-auditfslv
                       50G  231M   47G   1% /oradbaudit
/dev/mapper/datavg-dbtoolslv
                      9.8G   86M  9.2G   1% /oratools
/dev/asm/ggatevol-387
                       20G  562M   20G   3% /gg/GG11
                                                           
Start again………..

[root@oradev11 bin]# ./crsctl start res ora.crf -init
CRS-2672: Attempting to start 'ora.crf' on 'oradev11'

CRS-2676: Start of 'ora.crf' on 'oradev11' succeeded

[root@oradev11 bin]# ./crsctl status res -t -init
--------------------------------------------------------------------------------
NAME           TARGET  STATE        SERVER                   STATE_DETAILS
--------------------------------------------------------------------------------
Cluster Resources
--------------------------------------------------------------------------------
ora.asm
      1        ONLINE  ONLINE       oradev11           Started
ora.cluster_interconnect.haip
      1        ONLINE  ONLINE       oradev11
ora.crf
      1        ONLINE  ONLINE       oradev11
ora.crsd
      1        ONLINE  ONLINE       oradev11
ora.cssd
      1        ONLINE  ONLINE       oradev11
ora.cssdmonitor
      1        ONLINE  ONLINE       oradev11
ora.ctssd
      1        ONLINE  ONLINE       oradev11           OBSERVER
ora.diskmon
      1        OFFLINE OFFLINE
ora.drivers.acfs
      1        ONLINE  ONLINE       oradev11
ora.evmd
      1        ONLINE  ONLINE       oradev11
ora.gipcd
      1        ONLINE  ONLINE       oradev11
ora.gpnpd
      1        ONLINE  ONLINE       oradev11
ora.mdnsd
      1        ONLINE  ONLINE       oradev11


Now I see all the resources are up and running

Refer:
Oracle Cluster Health Monitor (CHM) using large amount of space (more than default) (Doc ID 1343105.1)



Friday, August 12, 2011

Single NODE RAC


What is Oracle RAC One Node?

Oracle introduced a new option called RAC One Node with the release of 11gR2 in late 2009. This option is available with Enterprise edition only. Basically, it provides a cold failover solution for Oracle databases. It’s a single instance of Oracle RAC running on one node of the cluster while the 2nd node is in a cold standby mode. If the instance fails for some reason, then RAC One Node detects it and first tries to restart the instance on the same node. The instance is relocated to the 2nd node in case there is a failure or fault in 1st node and the instance cannot be restarted on the same node. The benefit of this feature is that it automates the instance relocation without any downtime and does not need a manual intervention. It uses a technology called Omotion, which facilitates the instance migration/relocation. “RAC one” is Oracle’s answer or solution to OS clustering solution like Veritas Storage Foundation, Sun Solaris cluster, IBM HACMP, and HP Service guard etc.

Purpose

Its Oracle’s attempt to tie customers to a single vendor by eliminating the need to buy 3rd party OS cluster solutions. First, it introduced Oracle Clusterware with 10g and stopped the need to rely on 3rd party cluster software and now it intends to conquer the rest who are still using HACMP, Sun Solaris cluster etc. for cold failover.

Benefits

The Oracle RAC One node provides the following benefits:

•         Built-in cluster failover for high availability
•         Rolling patches for single instance database
•         Proactive migration / failover of the instance
•         Live migration of instances across servers
•         Online upgrade to RAC

The rolling upgrade is really useful. Upgrade to the OS, and Database can be done without any downtime unless upgrade requires some scripts to be run against the database. With RAC One Node, the DBA’s and Sys admins can be proactive and migrate/failover the instance to another node to perform any critical maintenance activity.

What it's not suited for

According to me the RAC one node is not a viable or recommended solution in the following scenarios:

•         To load balance unlike regular RAC
•         A true high availability solution
•         As a DR solution; Data guard best suits the bill
•         For mission critical applications

Cost

It is definitely not FREE. Oracle has priced RAC one at par with Active Data Guard. The RAC One node is priced separately and costs $10,000 per processor as against $23,000 for regular RAC. The licensing cost is required for ONE node only (in a 2-node setup). RAC one node is eligible for 10-day rule, allowing a customer to migrate to another without the need to buy additional license up to 10-days in a calendar year. People arguing against paying a license fee for resources they are not using will still lament.

Conclusion

I am still not very convinced on the usefulness of RAC one node. I think customers invest in RAC for their mission critical applications and achieving high availability and load balancing at the same time. Those who don’t go for RAC rely on Data Guard and now with 11g, on Active Data Guard. So don’t see a huge requirement for RAC One except seamless failover within a data center. The licensing is a bit disappointing; they are making clients pay $10 K. Moreover RAC is free with Standard edition though one doesn’t get enterprise features and limited to 4 CPU sockets only. So, thinking RAC One will be popular among customers who are currently using standard edition and want to switch to enterprise will be wrong. However, this is still a very new feature and as more people adopt it, we will get more clarity on its’ usability. I am planning to do a POC on it and would publish the installation steps and any findings (goods things and not so good things) of my POC.

Sunday, August 7, 2011

SINGLE CLIENT ACCESS NAME (SCAN)


Single Client Access Name (SCAN) is a new Oracle Real Application Clusters (RAC) 11g Release 2 feature that provides a single name for clients to access Oracle Databases running in a cluster. The benefit is that the client’s connect information does not need to change if you add or remove nodes in the cluster. Having a single name to access the cluster allows clients to use the EZConnect client and the simple JDBC thin URL to access any database running in the cluster, independently of which server(s) in the cluster the database is active. SCAN provides load balancing and failover for client connections to the database. The SCAN works as a cluster alias for databases in the cluster.


How does SCAN work ?




NETWORK REQUIREMENTS FOR USING SCAN

The SCAN is configured during the installation of Oracle Grid Infrastructure that is distributed with Oracle Database 11g Release2. Oracle Grid Infrastructure is a single Oracle Home that contains Oracle Clusterware and Oracle Automatic Storage Management. You must install Oracle Grid Infrastructure first in order to use Oracle RAC 11g Release 2. During the interview phase of the Oracle Grid Infrastructure installation, you will be prompted to provide a SCAN name. There are 2 options for defining the SCAN:

1. Define the SCAN in your corporate DNS (Domain Name Service)

2. Use the Grid Naming Service (GNS)


USING OPTION 1 - DEFINE THE SCAN IN YOUR CORPORATE DNS

If you choose Option 1, you must ask your network administrator to create a single name that resolves to 3 IP addresses using a round-robin algorithm. Three IP addresses are recommended considering load balancing and high availability requirements regardless of the number of servers in the cluster. The IP addresses must be on the same subnet as your public network in the cluster. The name must be 15 characters or less in length, not including the domain, and must be resolvable without the domain suffix (for example: “sales1-scan’ must be resolvable as opposed to “scan1-scan.example.com”). The IPs must not be assigned to a network interface (on the cluster), since Oracle Clusterware will take care of it.


You can check the SCAN configuration in DNS using “nslookup”. If your DNS is set up to provide round-robin access to the IPs resolved by the SCAN entry, then run the “nslookup” command at least twice to see the round-robin algorithm work. The result should be that each time, the “nslookup” would return a set of 3 IPs in a different order. 
Note: If your DNS server does not return a set of 3 IPs as shown in figure 3 or does not round-robin, ask your network administrator to enable such a setup. DNS using a round-robin algorithm on its own does not ensure failover of connections. However, the Oracle Client typically handles this. It is therefore recommended that the minimum version of the client used is the Oracle Database 11g Release 2 client.


USING OPTION 2 - THE GRID NAMING SERVICE (GNS)

If you choose option 2, you only need to enter the SCAN during the interview. During the cluster configuration, three IP addresses will be acquired from a DHCP service (using GNS assumes you have a DHCP service available on your public network) to create the SCAN and name resolution for the SCAN will be provided by the GNS1.


IF YOU DO NOT HAVE A DNS SERVER AVAILABLE AT INSTALLATION TIME

Oracle Universal Installer (OUI) enforces providing a SCAN resolution during the Oracle Grid Infrastructure installation, since the SCAN concept is an essential part during the creation of Oracle RAC 11g Release 2 databases in the cluster. All Oracle Database 11g Release 2 tools used to create a database (e.g. the Database Configuration Assistant (DBCA), or the Network Configuration Assistant (NetCA)) would assume its presence. Hence, OUI will not let you continue with the installation until you have provided a suitable SCAN resolution.

               However, in order to overcome the installation requirement without setting up a DNS-based SCAN resolution, you can use a hosts-file based workaround. In this case, you would use a typical hosts-file entry to resolve the SCAN to only 1 IP address and one IP address only. It is not possible to simulate the round-robin resolution that the DNS server does using a local host file. The host file look-up the OS performs will only return the first IP address that matches the name. Neither will you be able to do so in one entry (one line in the hosts-file). Thus, you will create only 1 SCAN for the cluster. (Note that you will have to change the hosts-file on all nodes in the cluster for this purpose.) This workaround might also be used when performing an upgrade from former (pre-Oracle Database 11g Release 2) releases. However, it is strongly recommended to enable the SCAN configuration as described under “Option 1” or “Option 2” above shortly after the upgrade or the initial installation. In order to make the cluster aware of the modified SCAN configuration, delete the entry in the hosts-file and then issue: "srvctl modify scan -n <scan_name>" as the root user on one node in the cluster. The scan_name provided can be the existing fully qualified name (or a new name), but should be resolved through DNS, having 3 IPs associated with it, as discussed. The remaining reconfiguration is then performed automatically.


SCAN CONFIGURATION IN THE CLUSTER

During cluster configuration, several resources are created in the cluster for SCAN. For each of the 3 IP addresses that the SCAN resolves to, a SCAN VIP resource is created and a SCAN Listener is created. The SCAN Listener is dependent on the SCAN VIP and the 3 SCAN VIPs (along with their associated listeners) will be dispersed across the cluster. This means, each pair of resources (SCAN VIP + Listener) will be started on a different server in the cluster, assuming the cluster consists of three or more nodes.

In case, a 2-node-cluster is used (for which 3 IPs are still recommended for simplification reasons), one server in the cluster will host two sets of SCAN resources under normal operations. If the node where a SCAN VIP is running fails, the SCAN VIP and its associated listener will failover to another node in the cluster. If by means of such a failure the number of available servers in the cluster becomes less than three, one server would again host two sets of SCAN resources. If a node becomes available in the cluster again, the formerly mentioned dispersion will take effect and relocate one set accordingly.


DATABASE CONFIGURATION USING SCAN

For Oracle Database 11g Release 2, SCAN is an essential part of the configuration and therefore the REMOTE_LISTENER parameter is set to the SCAN per default, assuming that the database is created using standard Oracle tools (e.g. the formerly mentioned DBCA). This allows the instances to register with the SCAN Listeners as remote listeners to provide information on what services are being provided by the instance, the current load, and a recommendation on how many incoming connections should be directed to the instance. In this context, the LOCAL_LISTENER parameter must be considered. The LOCAL_LISTENER parameter should be set to the node-VIP. If you need fully qualified domain names, ensure that LOCAL_LISTENER is set to the fully qualified domain name (e.g. node-VIP.example.com). By default, a node listener is created on each node in the cluster during cluster configuration. With Oracle Grid Infrastructure 11g Release 2 the node listener run out of the Oracle Grid Infrastructure home and listen on the node-VIP using the specified port (default port is 1521). Unlike in former database versions, it is not recommended to set your REMOTE_LISTENER parameter to a server side TNSNAMES alias that resolves the host to the SCAN (HOST=sales1-scan for example) in the address list entry, but use the simplified “SCAN:port” syntax as shown in figure 5.


 [oracle@mynode] srvctl config scan_listener
SCAN Listener LISTENER_SCAN1 exists. Port: TCP:1521

SCAN Listener LISTENER_SCAN2 exists. Port: TCP:1521

SCAN Listener LISTENER_SCAN3 exists. Port: TCP:1521


[oracle@mynode] srvctl config scan
SCAN name: sales1-scan, Network: 1/133.22.67.0/255.255.255.0/

SCAN VIP name: scan1, IP: /sales1-scan.example.com/133.22.67.192

SCAN VIP name: scan2, IP: /sales1-scan.example.com/133.22.67.193

SCAN VIP name: scan3, IP: /sales1-scan.example.com/133.22.67.194


NAME                  TYPE           VALUE

-----------             -----------    ---------------------

local_listener     string           (DESCRIPTION=

                                                     (ADDRESS_LIST=

                                         (ADDRESS=(PROTOCOL=TCP)(HOST=133.22.67.111)(PORT=1521))))

remote_listener    string             sales1-scan.example.com:1521


Note: if you are using the easy connect naming method, you may need to modify your SQLNET.ORA to ensure that EZCONNECT is in the list when specifying the order of the naming methods used for the client name resolution lookups (the Oracle 11g Release 2 default is NAMES.DIRECTORY_PATH=(tnsnames, ldap, ezconnect))


HOW CONNECTION LOAD BALANCING WORKS USING SCAN

For clients connecting using Oracle SQL*Net 11g Release 2, three IP addresses will be received by the client by resolving the SCAN name through DNS as discussed. The client will then go through the list it receives from the DNS and try connecting through one of the IPs received. If the client receives an error, it will try the other addresses before returning an error to the user or application. This is similar to how client connection failover works in previous releases when an address list is provided in the client connection string.

When a SCAN Listener receives a connection request, the SCAN Listener will check for the least loaded instance providing the requested service. It will then re-direct the connection request to the local listener on the node where the least loaded instance is running. Subsequently, the client will be given the address of the local listener. The local listener will finally create the connection to the database instance.


Note:  This example assumes an Oracle 11g R2 client using a default TNSNAMES. ORA:

ORCLservice =

(DESCRIPTION =

(ADDRESS = (PROTOCOL = TCP)(HOST = sales1-scan.example.com)(PORT = 1521))

(CONNECT_DATA =

(SERVER = DEDICATED)

(SERVICE_NAME = MyORCLservice)

))


VERSION AND BACKWARD COMPATIBILITY

The successful use of SCAN to connect to an Oracle RAC database in the cluster depends on the ability of the client to understand and use the SCAN as well as on the correct configuration of the REMOTE_LISTENER parameter setting in the database. If the version of the Oracle Client connecting to the database as well as the Oracle Database version used are both Oracle Database 11g Release 2 and the default configuration is used as described in this paper, no changes to the system are typically required.

The same holds true, if the Oracle Client version and the version of the Oracle Database that this client is connecting to are both pre-11g Release 2 version (e.g. Oracle Database 11g Release 1 or Oracle Database 10g Release 2, or older). In this case, the pre-11g Release 2 client would use a TNS connect descriptor that resolves to the node-VIPs of the cluster, while the Oracle pre-11g Release 2 database would still use a REMOTE_LISTENER entry pointing to the node-VIPs. The disadvantage of this configuration is that SCAN would not be used and hence the clients are still exposed to changes every time the cluster changes in the backend. Similarly, if an Oracle Database 11g Release 2 is used, but the clients remain on a former version. The solution is to change the Oracle client and / or Oracle Database REMOTE_LISTENER settings accordingly. The following cases need to be considered:


Note: If using a pre-11g Release 2 client (Oracle Database 11g Release or Oracle Database 10g Release 2, or older) you will not fully benefit from the advantages of SCAN. Reason: The Oracle Client will not be able to handle a set of three IPs returned by the DNS for SCAN. Hence, it will try to connect to only the first address returned in the list and will more or less ignore the others. If the SCAN Listener listening on this specific IP is not available or the IP itself is not available, the connection will fail. In order to ensure load balancing and connection failover with pre-11g Release 2 clients, you will need to change the TNSNAMES.ora of the client so that it would use 3 address lines, where each address line resolves to one of the SCAN VIPs.

    

Sample TNSNAMES.ora for Oracle Database pre- 11g Release 2 Clients

sales.example.com =(DESCRIPTION=

(ADDRESS_LIST= (LOAD_BALANCE=on)(FAILOVER=ON)

(ADDRESS=(PROTOCOL=tcp)(HOST=133.22.67.192)(PORT=1521))

(ADDRESS=(PROTOCOL=tcp)(HOST=133.22.67.193)(PORT=1521))

(ADDRESS=(PROTOCOL=tcp)(HOST=133.22.67.194)(PORT=1521)))

(CONNECT_DATA=(SERVICE_NAME= salesservice.example.com)))


USING SCAN IN A MAXIMUM AVAILABILITY ARCHITECTURE ENVIRONMENT

If you have implemented a Maximum Availability Architecture (MAA) environment, in which you use Oracle RAC for both your primary and standby database (in both, your primary and standby site), which are synchronized using Oracle Data Guard, using SCAN provides a simplified TNSNAMES configuration that a client can use to connect to the database independently of whether the primary or standby database is the currently active (primary) database. In order to use this simplified configuration, Oracle Database 11g Release 2 introduces two new SQL*Net parameters that can be used on for connection strings of individual clients. The first parameter is CONNECT_TIMEOUT. It specifies the timeout duration (in seconds) for a client to establish an Oracle Net connection to an Oracle database. This parameter overrides SQLNET.OUTBOUT_CONNECT_TIMEOUT in the SQLNET.ORA. The second parameter is RETRY_COUNT and it specifies the number of times an ADDRESS_LIST is traversed before the connection attempt is terminated. Using these two parameters, both, the SCAN on the primary site and the standby site, can be used in the client connection strings. Even, if the randomly selected address points to the site that is not currently active, the timeout will allow the connection request to failover before the client has waited unreasonably long (the default timeout depending on the operating system can be as long as 10 minutes).


sales.example.com =(DESCRIPTION= (CONNECT_TIMEOUT=10)(RETRY_COUNT=3)
                              (ADDRESS_LIST= (LOAD_BALANCE=on)(FAILOVER=ON)

                             (ADDRESS=(PROTOCOL=tcp)(HOST=sales1-scan)(PORT=1521))

                             (ADDRESS=(PROTOCOL=tcp)(HOST=sales2-scan)(PORT=1521)))

                             (CONNECT_DATA=(SERVICE_NAME= salesservice.example.com)))



Sunday, July 24, 2011

Apply CRS Patch to CRS_HOME only or RDBMS_HOME only


Goal

The Clusterware patches for 10g and 11g contain binaries that can be installed in the Clusterware home (CRS_HOME), and the database home (RDBMS_HOME). Due to normal upgrading of the Oracle software in a RAC environment, customers may want to maintain multiple Oracle homes of different versions. Under some situation, customers may need to apply a CRS patch to either CRS_HOME or RDBMS_HOME.

The README.txt that comes with a CRS patch generally assumes customers apply the patch to both CRS_HOME and RDBMS_HOME at the same time. This note gives instructions on what to do if customers need to apply a CRS patch to only CRS_HOME or only RDBMS_HOME.

Customers may need to do this when:

1). Customers have a mixed CRS and RDBMS version.
2). Customers need to apply a CRS patch to a single ASM instance environment, or single instance Exadata ennvironment. For example, patch ocssd.bin on single ASM instance environment; patch diskmon.bin on a single instance Exadata environment.

Solution

1. Apply a CRS patch to CRS_HOME but not to RDBMS_HOME.

Under this scenario, since CRS patch is not applied to RDBMS_HOME, one can skip the steps in README.txt related to RDBMS_HOME. An outline of the steps follows.
1). Verify that the Oracle Inventory is properly configured.
2). Unzip the PSE container file
3). Shutdown the RDBMS and ASM instances, listeners and nodeapps on all nodes before shutting down the CRS daemons.
4). Invoke the custom/scripts/prerootpatch.sh as root to unlock protected files.
5). Invoke custom/scripts/prepatch.sh -crshome <CRS_HOME>
     Note: Do not run custom/server/7<patch#>/custom/scripts/prepatch.sh -dbhome <RDBMS_HOME>
6). Patch the CRS_HOME files
     opatch napply -local -oh <CRS_HOME> -id <patch#>
     Note: Do not run opatch against RDBMS_HOME.
7). Configure CRS_HOME
     custom/scripts/postpatch.sh -crshome <CRS_HOME>
     Note: Do not run custom/server/<patch#>/custom/scripts/postpatch.sh -dbhome <RDBMS_HOME>
8). Invoke custom/scripts/postrootpatch.sh -crshome <CRS_HOME> as root
9). Determine whether the patch has been installed on CRS_HOME:
     opatch lsinventory -detail -oh <CRS_HOME>
     Note: Since the patch is not installed on RDBMS_HOME, opatch lsinventory -oh <RDBMS_HOME> will not show the patch#.

2. Apply CRS patch on RDBMS_HOME but not CRS_HOME.

Under this scenario, since CRS patch is not applied to CRS_HOME, but it is applied to RDBMS_HOME, one can skip the steps in README.txt related to CRS_HOME.
1). Verify that the Oracle Inventory is properly configured.
2). Unzip the PSE container file
3). Shutdown the RAC/ASM instances, listeners, nodeapps that runs from RDBMS_HOME. Since we are not patching CRS_HOME, CRS daemons do not need to be down.
4). Skip:  custom/scripts/prerootpatch.sh -crshome <CRS_HOME> -crsuser <username>
5). Skip: custom/scripts/prepatch.sh -crshome <CRS_HOME>
     Run: custom/server/<patch#>/custom/scripts/prepatch.sh -dbhome <RDBMS_HOME>
6). Patch the Files
     Skip 6.1 Patch the CRS home files:
                 opatch napply -local -oh <CRS_HOME> -id <patch#>
     Run: 6.2 Patch the RDBMS home files:
                 opatch napply custom/server/ -local -oh <RDBMS_HOME> -id <patch#>
7). Configure the HOME
     Skip: 7.1 Configure the CRS HOME
             custom/scripts/postpatch.sh -crshome <CRS_HOME>
     Run: 7.2 Configure the RDBMS HOME
             custom/server/<patch#>/custom/scripts/postpatch.sh -dbhome <RDBMS_HOME>
8). Skip: custom/scripts/postrootpatch.sh -crshome <CRS_HOME>
9). Verify patch is installed on RDBMS_HOME:
      opatch lsinventory -detail -oh <RDBMS_HOME>
10). Start nodeapps, listeners, ASM/RAC instances that were shutdown before the patching

Auto Scroll Stop Scroll