Lapsed, fee not paid5 drawingsSystem and method for detecting and predicting anomalies based on analysis of time-series data
Provided are a system and method for detecting and predicting anomalies based on analysis of time-series data.
US 9,952,936 B2 · Assignee: Hitachi, Ltd. · Inventors: Hayasaka; Mitsuo et al.
Sheet 1 of 18 from the published document. All sheets in the USPTO PDF
In a storage system for backing up data of an external apparatus, the external apparatus and a storage apparatus collaboratively perform efficient de-duplication. A storage system stores data from the external apparatus in a unit of content, and includes a backup apparatus configured to execute backup processing to create backup data of the data from the external apparatus in the unit of content; and a storage apparatus coupled to the backup apparatus in a communication-enabled manner and configured to store the backup data received from the backup apparatus. A first backup processing part of the backup apparatus determines whether or not a content is already stored in the storage apparatus by using first redundancy determination information that is information for determining whether or not each of contents of the backup data is already stored in the storage apparatus.
A storage apparatus is coupled to an external apparatus such as a host computer via a communication network. The storage apparatus of this type includes, for example, multiple hard disk drives (HDD) as storage devices for storing data. For the purpose of reducing costs required for storage media, data volume reduction processing is performed in the course of storing data into a storage device. The reduction of data volume is done through file compression or de-duplication. The file compression reduces the data volume by compressing data segments having identical data within a single file. On the other hand, the de-duplication reduces the total data volume in a file system or a storage system by detecting data segments having identical data not only within a single file but also among files and by compressing the detected data segments. In the following description, a “chunk” denotes a un
1 of 18 drawing sheets so far from the published document, cropped to the drawing. Every sheet is in the USPTO PDF.
What the patent claimed, word for word. All of it is now free to use.
The present invention relates to a storage system for storing data from an external apparatus and a method of controlling the storage system.
A storage apparatus is coupled to an external apparatus such as a host computer via a communication network. The storage apparatus of this type includes, for example, multiple hard disk drives (HDD) as storage devices for storing data. For the purpose of reducing costs required for storage media, data volume reduction processing is performed in the course of storing data into a storage device. The reduction of data volume is done through file compression or de-duplication. The file compression reduces the data volume by compressing data segments having identical data within a single file. On the other hand, the de-duplication reduces the total data volume in a file system or a storage system by detecting data segments having identical data not only within a single file but also among files and by compressing the detected data segments.
In the following description, a “chunk” denotes a unit-data segment that is a unit for de-duplication. In addition, a “container” denotes a set of data consisting of two or more chunks. In general, chunks closely related to each other are gathered in a container for efficient execution of the de-duplication. Moreover, a “container index” denotes a table created for each container and including records of hash values calculated for the respective chunks stored in the container. In addition, a “content” denotes a logically-grouped set of data that is a unit for storage in the storage device. Types of contents include not only general files, but also an aggregate of general files, such as an archive file, a backup file and a virtual volume file.
For storing a content into a storage apparatus from a host computer via a communication network, there have been proposed techniques of performing efficient data storage processing through reduction in a data transfer volume by making redundancy determination in advance as to whether each chunk is already stored in the storage apparatus and by sending only chunks determined as not stored. These techniques are described in Patent Literatures 1 and 2, for example. CITATION LIST Patent Literature
[PTL 1] U.S. Pat. No. 5,990,180
[PTL 2] International Patent Application Publication No. WO2012/101674 SUMMARY OF INVENTION Technical Problem
For the redundancy determination of chunks, there are two determination methods; one is done by a host computer, and the other is done by a storage apparatus.
For example, above Patent Literature 1 describes a method in which the storage apparatus makes the redundancy determination, and the host computer acquires the result of the redundancy determination and stores only new chunks into the storage apparatus.
In the de-duplication method in Patent Literature 1, however, the host computer requests the storage apparatus to make redundancy determination for all the chunks. To this end, the host computer needs to transmit information necessary for the redundancy determination to the storage apparatus and receive the redundancy determination results from the storage apparatus. In this case, the performance is lowered by an amount corresponding to a data roundtrip between the host computer and the storage apparatus, as compared with the method in which the host computer makes redundancy determination alone.
In contrast, Patent Literature 2 describes a method in which the host computer makes the redundancy determination and stores new chunks into the storage apparatus based on the determination results.
In the de-duplication method in Patent Literature 2, however, the host computer needs to store information for use in the redundancy determination for all the data stored from the host computer into the storage apparatus. In this de-duplication method, for example, as stored data increases, the information for use in the redundancy determination increases in size and accordingly places a larger burden on the disk capacity of the host computer, which makes the de-duplication inefficient in some cases.
In view of these circumstances, the present invention aims to provide a storage system and a method of controlling the storage system, which enable efficient de-duplication through collaboration between an external apparatus and a storage apparatus. Solution to Problem
In order to solve the forging problem and other problems, one aspect of the present invention provides a storage system to store data from an external apparatus in a unit of content, the storage system including a backup apparatus configured to execute backup processing to create backup data of the data from the external apparatus in the unit of content, and a storage apparatus coupled to the backup apparatus in a communication-enabled manner and configured to store the backup data received from the backup apparatus. The backup apparatus includes first redundancy determination information that is information for determining whether or not a content of the backup data is already stored in the storage apparatus, and a first backup processing part configured to determine whether or not the content is already stored in the storage apparatus by using the first redundancy determination information. The storage apparatus includes second redundancy determination information that is information for determining whether or not the content of the backup data is already stored in the storage apparatus, and a second backup processing part configured to determine whether or not the content is already stored in the storage apparatus by using the second redundancy determination information. When the first backup processing part determines that the content of the backup data is not stored in the storage apparatus but the second backup processing part determines that the content is stored in the storage apparatus, the second backup processing part sends the second redundancy determination information to the backup apparatus, and the first backup processing part of the backup apparatus incorporates the received second redundancy determination information into the first redundancy determination information. Advantageous Effects of Invention
According to the present invention, provided are a storage system and a method of controlling the storage system, which enable efficient de-duplication through collaboration between an external apparatus and a storage apparatus. Problems, configurations and effects other than those described above will be clarified through the following description of embodiments.
FIG. 1 is a diagram illustrating an overall configuration of a storage system 1 according to a first embodiment of the present invention.
FIG. 2 is a block diagram illustrating configurations of a backup server and a storage apparatus according to the first embodiment.
FIG. 3 is a diagram illustrating a configuration example of a container index and a chunk index used in backup processing of the storage system 1 .
FIG. 4 is a diagram illustrating a configuration example of a container index and a chunk index used in restore processing of the storage system 1 .
FIG. 5 is a diagram illustrating a configuration example of a content index used in the restore processing of the storage system 1 .
FIG. 6 is a diagram explaining a concept of backup processing according to the first embodiment.
FIG. 7 is a flowchart illustrating an example of a processing procedure in the backup processing according to the first embodiment.
FIG. 8 is a flowchart illustrating an example of a processing procedure in the restore processing according to the first embodiment.
FIG. 9 is a diagram illustrating an overall configuration of a storage system 1 according to a second embodiment of the present invention.
FIG. 10 is a flowchart illustrating an example of a processing procedure in backup processing according to the second embodiment.
FIG. 11 is a diagram illustrating an overall configuration of a storage system 1 according to a third embodiment of the present invention.
FIG. 12 is a block diagram illustrating configurations of a backup server and a storage apparatus according to the third embodiment.
FIG. 13 is a flowchart illustrating an example of a processing procedure in backup processing according to the third embodiment.
FIG. 14 is a diagram illustrating an overall configuration of a storage system 1 according to a fourth embodiment of the present invention.
FIG. 15 is a flowchart illustrating an example of a processing procedure in backup processing according to the fourth embodiment.
FIG. 16 is a flowchart illustrating an example of a processing procedure in backup processing according to a fifth embodiment.
FIG. 17 is a flowchart illustrating an example of a processing procedure in backup processing according to a sixth embodiment.
FIG. 18 is a flowchart illustrating an example of a processing procedure in backup processing according to a seventh embodiment.
Hereinafter, embodiments of the present invention are described based on the drawings. It should be noted that the present invention is not limited to the following embodiments, but can be altered variously within the technical idea of the present invention. First Embodiment
Configuration of Storage System in First Embodiment
FIG. 1 illustrates an overall configuration of a storage system 1 according to a first embodiment of the present invention. This storage system 1 includes backup servers 14 ( 14 a , 14 b , . . . , 14 n ) installed in multiple sites 2 ( 2 a , 2 b , . . . , 2 n ), respectively, and a storage apparatus 10 installed in a data center 3 . In this description, the signs a, b, . . . , n are omitted in some cases where general description for each site is provided without distinction among the sites.
The multiple sites 2 and the data center 3 are coupled to each other via a communication network 4 . The communication network 4 can be formed of an appropriate communication line including, for example, a WAN (wide area network), a LAN (local area network), the internet, a public line or a private line. Each of the sites 2 includes a business server 5 ( 5 a , 5 b , . . . , 5 n ), clients 6 ( 6 a , 6 b , . . . , 6 n ), and a backup server 14 ( 14 a , 14 b , . . . , 14 n ). The business server 5 , the clients 6 , and the backup server 14 are coupled to each other in a communication-enabled manner via a communication network 13 ( 13 a , 13 b , . . . , 13 n ) such as a LAN, for example.
The business server 5 is a computer configured to receive a request from each of the clients 6 , and provide a service corresponding to the request. The business server 5 includes a processor such as a CPU (Central Processing Unit), a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory), an auxiliary storage device (not illustrated) such as a HDD (Hard Disk Drive) or a SSD (Solid State Drive), and the like. The client 6 is also a computer having substantially the same configuration as the business server 5 , and functions as a terminal for a user who makes use of services provided by the business server 5 .
On a regular basis, the backup server 14 backs up data of the business server 5 and the clients 6 coupled to the backup server 14 in each of the sites 2 , and sends the backup data to the data center 3 . In addition, in response to a request from either of the business server 5 and the client 6 , the backup server 14 performs data restore from the data center 3 into the business server 5 or client 6 .
In the data center 3 , the storage apparatus 10 stores data received from the multiple backup servers 14 into a storage medium of the storage apparatus 10 . Moreover, in response to a request from any of the backup servers 14 , the storage apparatus 10 loads data stored in the storage medium and sends the data to the backup server 14 .
In the storage system 1 of this embodiment, the backup servers 14 installed in the respective sites 2 and the storage apparatus 10 installed in the data center 3 collaboratively perform data de-duplication efficiently.
FIG. 2 is a block diagram illustrating configuration examples of the backup server 14 a installed in the site 2 a and the storage apparatus 10 installed in the data center 3 in the storage system 1 illustrated in FIG. 1 . FIG. 2 illustrates only the backup server 14 a , but the other backup servers 14 b , . . . , 14 n installed in the sites 2 b to 2 n have substantially the same configuration.
To begin with, the configuration example of the backup server 14 a is described. As illustrated in FIG. 2 , the backup server 14 a mainly includes a processor 102 such as a CPU (Central Processing Unit), a memory 103 such as a RAM (Random Access Memory) or a ROM (Read Only Memory), and an auxiliary storage device 104 such as a HDD (Hard Disk Drive) or a SSD (Solid State Drive) (the auxiliary storage device 104 will be referred to as a “HDD” below), and a network interface 105 , such as a NIC (Network Interface Card), which is a communication interface with the communication network 4 . The processor 102 , the memory 103 , the HDD 104 , and the network interface 105 are coupled to each other in a communication-enabled manner via a system bus 108 .
The processor 102 functions as an arithmetic processing unit including the CPU and the like, and controls the operations of the backup server 14 a in accordance with programs, operational parameters and the like stored in the memory 103 .
The memory 103 stores a backup program 106 and a restore program 107 on the backup server 14 side. Moreover, the memory 103 is used to store various kinds of information loaded from the HDD 104 , and is also used as a work memory for the processor 102 . The HDD 104 stores various kinds of software, management information, backup data, and the like. Incidentally, the backup program 106 and the restore program 107 may be stored in the HDD 104 and are loaded from the HDD 104 to the memory 103 when the processor 102 executes these programs.
The memory 103 also stores a container index 320 (first redundancy determination information) that is a table to which the backup program 106 and the restore program 107 reference while running. Instead, the container indexes 320 may be rolled out to the HDD 104 and rolled in the memory 103 as needed when the backup program 106 and the restore program 107 reference thereto.
Next, description is provided for the programs executed by the backup server 14 a . The backup program 106 provides functions to determine backup data and perform data processing such as redundancy determination processing, and sends backup data to the storage apparatus 10 via the network interface 105 . Moreover, the backup program 106 receives information necessary for the redundancy determination processing from the storage apparatus 10 via the network interface 105 .
The restore program 107 receives backup data necessary for restore processing from the storage apparatus 10 via the network interface 105 and restores the backup data to original data.
Next, the configuration example of the storage apparatus 10 is described. As similar to the backup server 14 a , the storage apparatus 10 mainly includes a processor 112 , a memory 113 , a HDD 114 , and a network interface 115 . The processor 112 , the memory 113 , the HDD 114 , and the network interface 115 are coupled to each other in a communication-enabled manner via a system bus 118 .
The processor 112 functions as an arithmetic processing unit including the CPU and the like, and controls the operations of the storage apparatus 10 in accordance with programs, operational parameters and the like stored in the memory 113 .
The memory 113 stores a backup program 116 and a restore program 117 on the storage apparatus 10 side. Moreover, the memory 113 is used to store various kinds of information loaded from the HDD 114 , and is also used as a work memory for the processor 112 . The HDD 114 stores various kinds of software, management information, backup data, and the like. The HDD 114 stores various kinds of software, management information and data after de-duplication.
The memory 113 also stores a chunk index 310 , a container index 320 , and a content index 370 that are tables to which the backup program 116 and the restore program 117 reference while running. Instead, the chunk index 310 , the container index 320 and the content index 370 may be rolled out to the HDD 114 and rolled in the memory 113 as needed when the backup program 116 and the restore program 117 reference thereto. The chunk index 310 , the container index 320 and the content index 370 constitute second redundancy determination information.
Next, description is provided for the programs executed by the storage apparatus 10 . The backup program 116 performs de-duplication on data received from the backup server 14 a and stores the data after the de-duplication into the HDD 114 . The backup program 116 sends information necessary for the backup server 14 a to execute the redundancy determination processing to the backup server 14 a via the network interface 115 .
The restore program 117 loads from the HDD 114 data corresponding to a restore request from the backup server 14 a and sends the data to the backup server 14 a via the network interface 115 .
Outlines of Backup Processing and Restore Processing
Subsequently, description is provided for outlines of backup processing and restore processing according to this embodiment.
Outline of De-Duplication Function Installed in Storage Apparatus 10
First, the outline of a de-duplication function according to this embodiment is described. The backup program 106 of the backup server 14 and the backup program 116 of the storage apparatus 10 according to this embodiment are equipped with a processing function to reduce a data volume of backup target data. The reduction of data volume is done through data processing such for example as file compression and de-duplication. The file compression is processing to reduce the data volume by compressing data segments (unit-data segments) having identical data within a single file. On the other hand, the de-duplication is processing to reduce the total data volume of data stored in a file system, a storage system or the like by compressing data segments having identical data detected not only within a single file, but also among files.
In the description of this embodiment, a data segment as a unit for de-duplication on backup data is called a “chunk,” and a set of data consisting of multiple chunks is called a “container.” In addition, a logically-grouped set of data that is a unit for storage in a storage device is called a “content.” Types of contents include not only general files, but also a collection of general files, such as an archive file, a backup file and a virtual volume file.
A container is formed to include a set of chunks closely related to each other. For example, each container is set in advance to have a predetermined number of chunks or a predetermined data volume, and chunks generated from a single content or two or more contents are gathered in the container until the container is filled up. In this way, it is possible to form a container in consideration of the locality of data. In other words, in the case of restoring backup data of a particular content to the original data, if once a container which stores the first chunk of the content is identified, it is highly possible to acquire the following chunks from the same container. Hence, it is expected to reduce the processing of loading different containers from the HDD 114 to the memory 113 for the purpose of restoring a particular content.
In general, the size of a single chunk is several kilo bytes or more. For this reason, the execution of the redundancy determination processing requires a long processing time and high costs if the chunks themselves are compared with each other one by one from the first chunk. To avoid this, the storage apparatus 10 according to this embodiment uses message digests of chunks and thereby enables execution of the redundancy determination processing within a short time at low costs. The message digest is a technique of outputting a digest with a fixed length in response to an input of data with an arbitrary length. In this description, an output result of the message digest is referred to as a “finger print (FP).” The finger print can be acquired by using an arbitrary hash function. As this hash function, a preferably usable one is a hash function that is highly likely to determine hash values, such for example as SHA256, which have very high randomness and unique to respective chunks.
To begin with, description is provided for a de-duplication method for each chunk in the backup program 106 of the backup server 14 .
Before sending a certain chunk to the storage apparatus 10 , the backup program 106 of the backup server 14 determines whether or not the chunk to be sent is a chunk having the same data as that already stored in the storage apparatus 10 (hereinafter called a “redundant chunk”) or a chunk not stored yet (hereinafter called a “new chunk”). Here, since the backup server 14 does not have the information of all the chunks stored in the storage apparatus 10 , the backup server 14 may wrongly determine a redundant chunk as a new chunk in some cases.
Then, when determining that the chunk to be sent is a new chunk, the backup program 106 sends the chuck and the finger print (hash value) of the chunk to the storage apparatus 10 . When the backup server 14 determines the chunk as the new chunk, the storage apparatus 10 also performs de-duplication for the chunk as described later, and therefore the redundant chunk is not registered redundantly. On the other hand, when determining the chunk as a redundant chunk, the backup program 106 sends the storage apparatus 10 link information indicating the storage location of the chunk instead of sending the chuck to the storage apparatus 10 .
Next, description is provided for a de-duplication method for each chunk of the backup program 116 of the storage apparatus 10 .
Before storing a certain chunk received from the backup server 14 into the HDD 114 , the backup program 116 of the storage apparatus 10 determines whether or not the received chunk is a redundant chunk that is a chunk having the same data as that of a chunk already stored in the HDD 114 or a new chunk not stored yet in the HDD 114 .
When determining that the received chunk is a new chunk, the backup program 116 directly stores the chunk into the HDD 114 . On the other hand, when determining that the received chunk is a redundant chunk, the backup program 116 stores the link information indicating the storage location of the redundant chunk into the HDD 114 , instead of storing the chunk into the HDD 114 . Alternatively, when receiving the link information of the redundant chunk from the backup server 14 , the backup program 116 directly stores the link information into the HDD 114 .
In the foregoing way, the backup program 106 of the backup server 14 and the backup program 116 of the storage apparatus 10 according to this embodiment repeatedly execute the de-duplication of chunks in collaboration with each other to prevent the duplicate registration of redundant chunks. This de-duplication of redundant chunks leads to a reduction in a used capacity of the HDD 114 and speed-up of the backup processing.
As described above, a “container” is a transaction unit for data storage into the HDD 114 , and includes multiple chunks into which a single content or two or more contents are split. In addition, the backup program 116 of the storage apparatus 10 creates a container index for each container to manage the locations of the respective chunks included in the container. In the container index, the offset of a chunk (the location of the chunk in the container) and the size of the chunk are stored. The container index is used for redundancy determination of chunks.
In addition, the backup program 116 of the storage apparatus 10 also creates a chunk index. The chunk index is a table indicating which container index stores each of chunks generated by splitting backup data. The storage apparatus 10 creates the chunk index when determining a container where to store a chunk. During execution of the backup processing, the chunk index is used to determine which container index to use for the redundancy determination of a chunk. The details of the container index and the chunk index will be described later.
In this embodiment, the finger prints of the respective chunks are stored in the foregoing container index, and the finger prints of chunks are compared with each other in the redundancy determination processing. This comparison leads to speed-up and cost-down of the redundancy determination processing as compared with the case where chunks are compared with each other on a bit-by-bit basis.
Incidentally, this embodiment may use a write-once storage device for the purposes of ensuring the integrity of data and implementing highly-reliable backup processing. In the write-once storage device, data can be written only once, but the written data can be read any number of times. Since data written into a write-once storage device is neither erasable nor changeable, the device of this type is suitable to data archives for preservation of evidence. As the write-once storage device, there is an optical disc device using a ROM optical disc, for example. In general, a magnetic disk device is not a write-once storage device because data written therein can be updated. However, the magnetic disk device may be used as a write-once storage device if a file system, a driver unit and the like in the device are configured to only allow appending of data (in other words, prohibit overwriting of data). In this embodiment, it is preferable to principally use a write-once hard disk drive suitable for data backup as the storage device for backup.
The foregoing container is set in advance to have a predetermined number of chunks or a predetermined data volume. Thus, chunks are gathered on the memory 113 until the container is filled up, and the chunks for each container are written to the storage device (the HDD 114 ) for backup when the container is filled up. For example, when a write-once hard disk drive is used as the storage device, the storage apparatus 10 appends chunks in a container on the memory 113 until the container is filled up. At the same time, the storage apparatus 10 creates the container index for managing the locations of chunks in the container and the chunk index for managing correspondences between the chunks and the container indexes. Note that the backup data includes a universal chunk which always appears in every backup generation, and the universal chunk is stored in a container prepared at the initial backup processing.
Configurations of Various Indexes in this Embodiment
Next, description is provided for configuration examples of the chunk index 310 and the container index 320 in this embodiment. FIGS. 3 and 4 illustrate the configuration example of the container index 320 and the configuration example of the chunk index 310 used in the backup processing and the restore processing in this embodiment. The container index 320 is a table created on a container-by-container basis. The chunk index 310 is a table for managing chunks stored in containers.
FIG. 3 illustrates a container index Tg( 320 ) created for a particular container among the container indexes 320 . The container index Tg( 320 ) includes items of a finger print 321 , a container offset 322 , and a chunk length 323 .
The finger print 321 stores a finger print (a hash value calculated through an appropriate hash function in this embodiment) of each chunk. The container offset 322 stores an offset value giving the head position of each chunk in the container. The chunk length 323 stores information indicating a chunk length. In short, each row of the container index Tg ( 320 ) stores management information on a chunk. The container index Tg ( 320 ) illustrated in FIG. 3 stores management information 320 b on a chunk b, management information 320 c on a chunk c, and management information 320 f on a chunk f. A letter indicating each chunk is added as a suffix to the management information on the chunk. For example, the fingerprint 321 calculated for the chunk b is expressed as FPb.
Multiple container indexes 320 are managed by use of the chunk index 310 . In the chunk index 310 , a container ID 312 that is a code for distinguishing a container from other containers, and a fingerprint 311 of each chunk are recorded in association with each other. The container ID 312 herein is also used as pointer information based on which the container index 320 can be referenced to. In this embodiment, the container ID 312 and the corresponding container index 320 use a common identifier called a UUID (Universally Unique IDentifier).
Note that whether or not to reference to the chunk index 310 may be determined based on a processing result of filter processing for identifying whether or not a chunk is a new chunk. More specifically, for a chunk certainly unrecorded in the chunk index 310 , reference processing to the chunk index 310 may be skipped entirely and the chunk may be directly stored in a new container. If this processing method is employed, the number of times the backup program 116 of the storage apparatus 10 references to the chunk index 310 can be reduced, and accordingly the speed-up of the backup processing can be further enhanced.
Here, for example, let us assume that four files of the container 380 , the container index 320 , the chunk index 310 , and a content index 370 are stored under four directories, respectively, in the HDD 114 of the storage apparatus 10 .
Container/uuid-Cf: Container main body
ContainerIndex/uuid-Cf: Container index database
ChunkIndex/Most-significant N bits of fp: Chunk index database
Contents/uuid-Cf: Content index database
In the example in FIG. 3 , the finger prints 311 and the container IDs 312 of all the chunks are registered in the chunk index 310 , but the number of registered chunks may be reduced. As described above, each container 380 is formed in consideration of the locality of data. In addition, the backup data includes many data segments that are identical or only partially corrected among the backup generations. For this reason, if a chunk stored in a certain container 380 is included in a particular content, the other chunks in the same container are very highly likely to be stored in the same content. Hence, after the container index 320 is searched out for a certain chunk from the chunk index 310 , the redundancy determination on the following chunks can be made by using the container index 320 thus searched out. FIG. 4 illustrates an example of the container indexes 320 and the chunk index 310 in the case where the number of chunks registered in the chunk index 310 is reduced.
Here, let us consider the execution of the backup processing for a content including a chunk b, a chunk c, and a chunk f under the condition that the container index Tg( 320 ) is not expanded on the memory 113 of the storage apparatus 10 , for example. Firstly, the backup program 116 of the storage apparatus 10 searches the chunk index 310 by using the finger print FPb of the chunk b. In the case of FIG. 3 , the finger print FPb is associated with the container ID Tg( 320 ). The backup program 116 loads the container index Tg( 320 ) from the HDD 114 , and expands the container index Tg ( 320 ) on the memory 113 . Then, the backup program 116 can make the redundancy determination on the chunk c and the chunk f by using the expanded container index Tg ( 320 ).
When the number of chunks registered in the chunk index 310 is reduced as described above, the storage capacity and memory usage necessary for the de-duplication can be reduced. In addition, the reduction in the number of chunks registered in the chunk index 310 leads to the speed-up of searching for the finger print 311 corresponding to any chunk.
Next, with reference to FIG. 5 , description is provided for a configuration example of the content index 370 used during execution of the restore processing. The content index 370 is a table created for each content to manage chunks included in the content. The content index 370 includes a content ID 371 , a finger print 372 , a container ID 373 , a content offset 374 , and a chunk length 375 .
The content ID 371 stores information for distinguishing a content from the other contents. The finger print 372 stores a finger print of each chunk (a hash value calculated for each chunk by using an appropriate hash function). The container ID 373 stores identification information for distinguishing a container storing each chunk from the other containers. The content offset 374 stores information indicating the location of each chunk in the content. The chunk length 375 stores information indicating a length of each chunk.
FIG. 5 , for example, illustrates S.sub.f1( 370 ), S.sub.f2( 370 ), S.sub.f3( 370 ), and S.sub.fn( 370 ) as examples of the content indexes. For example, by using the information S.sub.f3( 370 ) corresponding to a content f.sub.3, it can be known that the content f.sub.3 can be reconstructed from chunks b, c, d, e, and f, and it can be also known at which location each of the chunks b to f is stored in the content f.sub.3 based on the content offset 374 and the chunk length 375 .
The content offset 374 and the chunk length 375 of a content for which the content index 370 is formed indicate logical allocation of each chunk in the content. Also, the container offset 322 and the chunk length 323 in the foregoing container index 320 ( FIG. 3 ) indicate a logical allocation of each chunk in each container.
During execution of the restore processing, the restore program 117 of the storage apparatus 10 acquires the container ID 373 of each chunk in reference to the content index 370 , and searches for the container index 320 by using the container ID 373 . Then, on the basis of the storage location information of each chunk stored in the container index 320 , the restore program 117 acquires the chunks from the container 380 loaded from the HDD 114 . Thereafter, the restore program 117 reconstructs the content of a restore target according to the logical allocations in the content index 370 .
Outline of De-Duplication Function Installed in Storage System 1
Next, description is provided for an outline of de-duplication implemented by the storage system 1 in this embodiment. FIG. 6 schematically illustrates the outline of the de-duplication implemented by the storage system 1 in this embodiment. Here, although FIG. 6 illustrates only the backup server 14 a as the backup server 14 provided to the storage system 1 , multiple backup servers 14 ( 14 a , 14 b , . . . , 14 n ) are coupled to the storage apparatus 10 through the communication network 4 as in the case of FIG. 1 .
As illustrated in FIG. 6 , a backup target content includes a chunk a, a chunk b, a chunk c, a chunk d, a chunk e, and a chunk f. In addition, the storage apparatus 10 stores the chunk index U( 310 ) and the container indexes Tg( 320 ) and Tc( 320 ).
Here, let us consider the case where the backup server 14 a executes the first backup processing and no container index 320 is stored in the memory 103 or the HDD 104 in the backup server 14 a.
Firstly, the backup program 106 of the backup server 14 a makes the redundancy determination on the first chunk a. Since neither the memory 103 nor the HDD 104 stores the container index 320 , the backup program 106 determines that the chunk a is a new chunk, and sends the chunk a and the finger print FPa of the chunk a to the storage apparatus 10 .
The backup program 116 of the storage apparatus 10 makes the redundancy determination on the received chunk a by using the chunk index U( 310 ). Before this process, if the container index 320 is already expanded on the memory 113 of the storage apparatus 10 , the backup program 116 may reference to and search the container index 320 to find out whether or not the chunk a is stored redundantly. In reference to the chunk index U( 310 ), the backup program 116 determines that the chunk a is already stored in the container Tg( 380 ), hence processes the chunk a as a redundant chunk and sends the container index Tg( 320 ) to the backup server 14 a.
The backup program 106 of the backup server 14 a expands the received container index Tg( 320 ) on the memory 103 , and uses the container index Tg( 320 ) in the determination processing for the following chunks to be subjected to the backup processing. Since the container index Tg( 320 ) stores the finger prints FPb, FPc, and FPd of the chunk b, the chunk c, and the chunk d, respectively, the backup program 106 determines that the chunk b, the chunk c, and the chunk d are redundant chunks.
However, since the finger print FPe of the chunk e is not registered in the container index Tg( 320 ), the backup program 106 determines that the chunk e is a new chunk and sends the chunk e and the finger print FPe thereof to the storage apparatus 10 .
The backup program 116 of the storage apparatus 10 makes the redundancy determination on the chunk e as is the case with the chunk a and sends the relevant container index Tc ( 320 ) to the backup server 14 a.
The backup program 106 of the backup server 14 a performs the redundancy determination processing on the following chunk f by using both of the received container index Tc( 320 ) and the container index Tg( 320 ) already expanded on the memory 103 .
Since the chunks are gathered in each container in consideration of the locality of data as described above, there is a high possibility that the container Tg( 380 ) storing the chunk a may store the following chunk b, chunk c, and chunk d, and therefore efficient de-duplication can be performed.
Here, when the backup program 106 of the backup server 14 performs the redundancy determination processing, the backup program 106 references to the finger prints 321 of at least one container index 320 except for the case where neither the memory 103 nor the HDD 104 stores the container index 320 . To this end, the backup program 106 needs to expand the container indexes 320 on the memory 103 . However, the capacity of the memory 103 is limited, the memory 103 has difficulty in leaving all the container indexes 320 for use by the backup program 106 always expanded on the memory 103 . For this reason, the backup server 14 makes an effective utilization of the memory 103 by rolling in the container index 320 from the HDD 104 to the memory 103 , and by rolling out the container index 320 from the memory 103 to the HDD 104 . Here, the rolled-out container index 320 may be deleted from the HDD 104 . In addition, in the redundancy determination by the backup program 116 in the storage apparatus 10 , the roll-in and roll-out processing is similarly performed on the memory 113 and the HDD 114 of the storage apparatus 10 .
In this embodiment, the backup program 106 of the backup server 14 and the backup program 116 of the storage apparatus 10 make the redundancy determination by comparing the fingerprints 321 of the chunks with each other, but instead may make the redundancy determination by comparing the chunks themselves with each other on a bit-by-bit basis in order to enhance the reliability of the redundancy determination. In this case, the backup program 116 of the storage apparatus 10 sends the backup server 14 the main body of the container 380 including the target chunk.
Detailed Operation of Backup Processing in this Embodiment
The description continues in the full USPTO document.
About 6,905 words. The USPTO PDF has it with every drawing.
Fees are due 3.5, 7.5 and 11.5 years after grant. This patent expired on April 24, 2026, so the fee marked "not paid" was the one that went unpaid.
STORAGE SYSTEM AND METHOD OF CONTROLLING STORAGE SYSTEM
Filed Dec 2012 · published Jul 2015Storage system and method of controlling storage system
Filed Dec 2012 · granted Apr 2018Earlier publications, parents and continuations. None of them can still be enforced, or this patent would not be listed.
Prior art cited by the examiner or applicant. Useful when you check your own idea for novelty.
Everything on this page comes from the documents linked above.