Patent Yard Sign in
Lapsed, fee not paid

Voice retrieval apparatus, voice retrieval method, and non-transitory recording medium

US 9,767,790 B2 · Assignee: CASIO COMPUTER CO., LTD. · Inventors: Tomita; Hiroki

USPTO PDF

Overview

Sheet 1 of 13 from the published document. All sheets in the USPTO PDF

Abstract From the patent

A voice retrieval apparatus executes processes of: converting a retrieval string into a phoneme string; obtaining, from a time length memory, a continuous time length for each phoneme contained in the converted phoneme string; deriving a plurality of time lengths corresponding to a plurality of utterance rates as candidate utterance time lengths of voices corresponding to the retrieval string based on the obtained continuous time length; specifying, for each of the plurality of time lengths, a plurality of likelihood obtainment segments having the derived time length within a time length of a retrieval sound signal; obtaining a likelihood showing a plausibility that the specified likelihood obtainment segment specified is a segment where the voices are uttered; and identifying, based on the obtained likelihood, for each of the specified likelihood obtainment segments, an estimation segment where utterance of the voices is estimated in the retrieval sound signal.

Why it's free to use

  • The USPTO Official Gazette of November 18, 2025 lists it as expired on September 19, 2025 for an unpaid maintenance fee.
  • It isn't on any reinstatement notice published since.
  • Its 1 US relative has also lapsed, expired or never issued.
  • We check US rights only. Check foreign counterparts before selling abroad.
FiledNovember 30, 2015
GrantedSeptember 19, 2017
Expired (fee)September 19, 2025
Application number14/953729
Classification (CPC)G10L15/05 +3 more
Length18 claims · 29 pages

Background From the patent

Due to widespread popularization of multimedia contents, such as voice and motion image, there is a demand for a highly precise multimedia retrieval technology. With respect to such a technology, a voice retrieval technology that identifies a portion where voices corresponding to a retrieval term (query) subjected to retrieval is uttered in a sound signal has been studied. As for voice retrieval, a retrieval scheme with a sufficient performance has not been established yet in comparison with character string retrieval technologies based on image recognition. Hence, various technologies have been studied in order to realize a voice retrieval with a sufficient performance. For example, Non-patent Literature 1 (Y. Zhang and J. Glass, “An inner-product lower-bound estimate for dynamic time warping”, in Proc., ICASSP, 2011, pp. 5660-5663) discloses a method of comparing sound signals with eac

Drawings 13

8 of 13 drawing sheets so far from the published document, cropped to the drawing. Every sheet is in the USPTO PDF.

Figures as described

  • FIG. 1 is a diagram illustrating a physical structure of a voice retrieval apparatus according to a first embodiment of the present disclosure
  • FIG. 2 is a diagram illustrating a functional structure of the voice retrieval apparatus according to the first embodiment of the present disclosure
  • FIG. 3 is a diagram to explain a phoneme state
  • FIG. 4 is a diagram to explain how to derive an utterance time length corresponding to an utterance rate
  • FIG. 5A is a waveform diagram of a sound signal subjected to retrieval, FIG. 5B is a diagram illustrating a frame set in the sound signal subjected to retrieval, and FIG
  • FIG. 6 is a diagram illustrating an example way of performing Lower-Bounding on an output probability
  • FIG. 7 is a diagram to explain a selection method of a selector for candidate segments
  • FIG. 9 is a flowchart illustrating a flow of process of identifying a segment corresponding to a retrieval string
  • FIG. 10 is a diagram illustrating a functional structure of a voice retrieval apparatus according to a second embodiment of the present disclosure
  • FIG. 11A is a diagram to explain a method of a selector to select candidate segments after an obtained likelihood weight coefficient is multiplied, and FIG
  • FIG. 12 is a diagram to explain a selection method of the selector for candidate segments
  • FIG. 13B is a diagram illustrating an example comparison on a likelihood ranking corresponding to an utterance rate for each segment obtained by dividing the sound signal

Claims 18 total, 6 independent

What the patent claimed, word for word. All of it is now free to use.

  1. 1
    Independent claimA voice retrieval apparatus comprising: a processor; a memory that records a sound signal subjected to retrieval; and an output device comprising a screen, wherein the processor executes following processes: a converting process of converting a retrieval string into a phoneme string; a time length obtaining process of obtaining, from a database that stores continuous time length data on a phoneme, a continuous time length for each phoneme contained in the phoneme string converted in the converting process; a time length deriving process of deriving, based on the continuous time length obtained in the time length obtaining process, a plurality of time lengths corresponding to a plurality of utterance rates different from one another as candidate utterance time lengths of voices corresponding to the retrieval string; a segment specifying process of specifying, for each of the plurality of time lengths, a plurality of likelihood obtainment segments within the sound signal subjected to retrieval, each of the plurality of likelihood obtainment segments having the time length derived in the time length deriving process; a likelihood obtaining process of obtaining a likelihood showing a plausibility that the likelihood obtainment segment specified in the segment specifying process is a segment where voices corresponding to the retrieval string are uttered; an identifying process of identifying, based on the likelihood obtained in the likelihood obtaining process, an estimation segment where, within the sound signal subjected to retrieval, utterance of voices corresponding the retrieval string is estimated, the estimation segment being identified for each of the likelihood obtainment segments specified in the segment specifying process; and a displaying process of displaying on the screen of the output device, the estimation segment identified in the identifying process, wherein: the processor further executes a selecting process of selecting, based on the likelihood obtained in the likelihood obtaining process, one of the plurality of time lengths; and in the identifying process, the estimation segment is identified among the likelihood obtainment segments with the selected time length based on the likelihood obtained for the likelihood obtainment segment with the selected time length.
  2. 2
    The voice retrieval apparatus according to claim 1, wherein in the selecting process, an addition value obtained by adding together a predetermined number of the likelihoods obtained for the likelihood obtainment segments with a same time length in an order of higher likelihood is obtained for each of the plurality of time lengths, and the respective obtained addition values are compared one another, and the time length that has a maximum addition value is selected from the plurality of time lengths.
  3. 3
    The voice retrieval apparatus according to claim 2, wherein in the selecting process, the addition value is obtained by adding together the likelihoods obtained for the likelihood obtainment segments with the same time length, the likelihood to be added together being multiplied by a weight coefficient, wherein a higher the likelihood is, a larger the weight coefficient becomes.
  4. 4
    The voice retrieval apparatus according to claim 1, wherein: in the converting process, phonemes of an acoustic model that does not depend on adjacent phonemes are arranged in sequence to convert the retrieval string into the phoneme string; in the likelihood obtaining process, based on the phoneme string, the likelihood of the likelihood obtainment segment specified in the segment specifying process is obtained; in the selecting process, based on the likelihood obtained in the likelihood obtaining process, a plurality of candidates for the estimation segment is selected among the likelihood obtainment segment specified in the segment specifying process; the processor further executes: a second converting process of arranging phonemes of a second acoustic model that depends on adjacent phonemes in sequence, and converting the retrieval string into a second phoneme string; and a second likelihood obtaining process of obtaining, based on the second phoneme string for each of the plurality of candidates selected in the selecting process, a second likelihood showing a plausibility that the selected segment as the candidate of the estimation segment in the selecting process is a segment where voices corresponding to the retrieval string are uttered; and in the identifying process, based on the second likelihood obtained in the second likelihood obtaining process, the estimation segment is identified from the plurality of candidates selected in the selecting process.
  5. 5
    The voice retrieval apparatus according to claim 4, wherein in the selecting process, the plurality of candidates of the estimation segment is selected by selecting, from the likelihood obtainment segments beginning in the segment with a predetermined selection time length, the likelihood obtainment segment one by one with a maximum likelihood for each of the predetermined selection time lengths in the likelihood obtainment segments specified in the segment specifying process.
  6. 6
    Independent claimA voice retrieval apparatus comprising: a processor; a memory that records a sound signal subjected to retrieval; and an output device comprising a screen, wherein the processor executes following processes: a converting process of converting a retrieval string into a phoneme string; a time length obtaining process of obtaining, from a database that stores continuous time length data on a phoneme, a continuous time length for each phoneme contained in the phoneme string converted in the converting process; a time length deriving process of deriving, based on the continuous time length obtained in the time length obtaining process, a plurality of time lengths corresponding to a plurality of utterance rates different from one another as candidate utterance time lengths of voices corresponding to the retrieval string; a segment specifying process of specifying, for each of the plurality of time lengths, a plurality of likelihood obtainment segments within the sound signal subjected to retrieval, each of the plurality of likelihood obtainment segments having the time length derived in the time length deriving process; a likelihood obtaining process of obtaining a likelihood showing a plausibility that the likelihood obtainment segment specified in the segment specifying process is a segment where voices corresponding to the retrieval string are uttered; an identifying process of identifying, based on the likelihood obtained in the likelihood obtaining process, an estimation segment where, within the sound signal subjected to retrieval, utterance of voices corresponding the retrieval string is estimated, the estimation segment being identified for each of the likelihood obtainment segments specified in the segment specifying process; and a displaying process of displaying on the screen of the output device, the estimation segment identified in the identifying process, wherein: the processor further executes: a feature quantity calculating process of calculating a feature quantity of the sound signal subjected to retrieval in the likelihood obtainment segment specified in the segment specifying process for each frame that is a time window to compare the sound signal with an acoustic model; and an output probability obtaining process of obtaining, for each of the frames, an output probability that the feature quantity of the sound signal subjected to retrieval is output from each phoneme contained in the phoneme string, and in the likelihood obtaining process, respective values that are each a logarithm of the output probability obtained for each of the frames contained in the likelihood obtainment segment specified in the segment specifying process are added to obtain the likelihood of the likelihood obtainment segment.
  7. 7
    The voice retrieval apparatus according to claim 6, further comprising an output probability memory that stores, in association with each other for each of the frames contained in the sound signal subjected to retrieval, each phoneme state of an acoustic model, and, the output probability that the feature quantity of the sound signal subjected to retrieval is output from each phoneme state created from the acoustic model, wherein in the output probability obtaining process, when the retrieval string is converted into the phoneme string in the converting process, the output probability stored in association with each phoneme state contained in the phoneme string is obtained from the output probabilities stored in the output probability memory for each of the frames contained in the likelihood obtainment segment.
  8. 8
    The voice retrieval apparatus according to claim 7, wherein: the processor further executes a replacing process of replacing each of the output probabilities obtained in the output probability obtaining process for each of the frames with a maximum output probability in the frame, an N1 number of frames previous to the frame, and an N2 number of frames subsequent to the frame; the symbols N1 and N2 are each a natural number including zero, but either the N1 or the N2 is not zero; and in the likelihood obtaining process, based on the output probability having undergone replacement in the replacing process, the likelihood of the likelihood obtainment segment specified in the segment specifying process is obtained.
  9. 9
    Independent claimA voice retrieval method by a voice retrieval apparatus comprising a memory that records a sound signal subjected to retrieval, the method comprising: a converting step of converting a retrieval string into a phoneme string; a time length obtaining step of obtaining, from a database that stores continuous time length data on a phoneme, a continuous time length for each phoneme contained in the phoneme string converted in the converting step; a time length deriving step of deriving, based on the continuous time length obtained in the time length obtaining step, a plurality of time lengths corresponding to a plurality of utterance rates different from one another as candidate utterance time lengths of voices corresponding to the retrieval string; a segment specifying step of specifying, for each of the plurality of time lengths, a plurality of likelihood obtainment segments within the sound signal subjected to retrieval, each of the plurality of likelihood obtainment segments having the time length derived in the time length deriving step; a likelihood obtaining step of obtaining a likelihood showing a plausibility that the likelihood obtainment segment specified in the segment specifying step is a segment where voices corresponding to the retrieval string are uttered; an identifying step of identifying, based on the likelihood obtained in the likelihood obtaining step, an estimation segment where, within the sound signal subjected to retrieval, utterance of voices corresponding the retrieval string is estimated, the estimation segment being identified for each of the likelihood obtainment segments specified in the segment specifying step; and a displaying step of displaying on a screen of an output device, the estimation segment identified in the identifying step, wherein: the voice retrieval method further comprises a selecting step of selecting, based on the likelihood obtained in the likelihood obtaining step, one of the plurality of time lengths; and in the identifying step, the estimation segment is identified among the likelihood obtainment segments with the selected time length based on the likelihood obtained for the likelihood obtainment segment with the selected time length.
  10. 10
    The voice retrieval method according to claim 9, wherein in the selecting step, an addition value obtained by adding together a predetermined number of the likelihoods obtained for the likelihood obtainment segments with a same time length in an order of higher likelihood is obtained for each of the plurality of time lengths, and the respective obtained addition values are compared one another, and the time length that has a maximum addition value is selected from the plurality of time lengths.
  11. 11
    The voice retrieval method according to claim 10, wherein in the selecting step, the addition value is obtained by adding together the likelihoods obtained for the likelihood obtainment segments with the same time length, the likelihood to be added together being multiplied by a weight coefficient, wherein a higher the likelihood is, a larger the weight coefficient becomes.
  12. 12
    The voice retrieval method according to claim 9, wherein: in the converting step, phonemes of an acoustic model that does not depend on adjacent phonemes are arranged in sequence to convert the retrieval string into the phoneme string; in the likelihood obtaining step, based on the phoneme string, the likelihood of the likelihood obtainment segment specified in the segment specifying step is obtained; in the selecting step, based on the likelihood obtained in the likelihood obtaining step, a plurality of candidates for the estimation segment is selected among the likelihood obtainment segment specified in the segment specifying step; the method further comprises: a second converting step of arranging phonemes of a second acoustic model that depends on adjacent phonemes in sequence, and converting the retrieval string into a second phoneme string; and a second likelihood obtaining step of obtaining, based on the second phoneme string for each of the plurality of candidates selected in the selecting step, a second likelihood showing a plausibility that the selected segment as the candidate of the estimation segment in the selecting step is a segment where voices corresponding to the retrieval string are uttered; and in the identifying step, based on the second likelihood obtained in the second likelihood obtaining step, the estimation segment is identified among the plurality of candidates selected in the selecting step.
  13. 13
    The voice retrieval method according to claim 12, wherein in the selecting step, the plurality of candidates of the estimation segment is selected by selecting, from the likelihood obtainment segments beginning in the segment with a predetermined selection time length, the likelihood obtainment segment one by one with a maximum likelihood for each of the predetermined selection time lengths in the likelihood obtainment segments specified in the segment specifying step.
  14. 14
    Independent claimA voice retrieval method by a voice retrieval apparatus comprising a memory that records a sound signal subjected to retrieval, the method comprising: a converting step of converting a retrieval string into a phoneme string; a time length obtaining step of obtaining, from a database that stores continuous time length data on a phoneme, a continuous time length for each phoneme contained in the phoneme string converted in the converting step; a time length deriving step of deriving, based on the continuous time length obtained in the time length obtaining step, a plurality of time lengths corresponding to a plurality of utterance rates different from one another as candidate utterance time lengths of voices corresponding to the retrieval string; a segment specifying step of specifying, for each of the plurality of time lengths, a plurality of likelihood obtainment segments within the sound signal subjected to retrieval, each of the plurality of likelihood obtainment segments having the time length derived in the time length deriving step; a likelihood obtaining step of obtaining a likelihood showing a plausibility that the likelihood obtainment segment specified in the segment specifying step is a segment where voices corresponding to the retrieval string are uttered; an identifying step of identifying, based on the likelihood obtained in the likelihood obtaining step, an estimation segment where, within the sound signal subjected to retrieval, utterance of voices corresponding the retrieval string is estimated, the estimation segment being identified for each of the likelihood obtainment segments specified in the segment specifying step; and a displaying step of displaying on a screen of an output device, the estimation segment identified in the identifying step, wherein: the voice retrieval method further comprises: a feature quantity calculating step of calculating a feature quantity of the sound signal subjected to retrieval in the likelihood obtainment segment specified in the segment specifying step for each frame that is a time window to compare the sound signal with an acoustic model; and an output probability obtaining step of obtaining, for each frame, an output probability that the feature quantity of the sound signal subjected to retrieval is output from each phoneme contained in the phoneme string, and in the likelihood obtaining step, respective values that are each a logarithm of the output probability obtained for each frame contained in the likelihood obtainment segment specified in the segment specifying step are added to obtain the likelihood of the likelihood obtainment segment.
  15. 15
    The voice retrieval method according to claim 14, wherein: the voice retrieval apparatus further comprises an output probability memory that stores, in association with each other for each of the frames contained in the sound signal subjected to retrieval, each phoneme state of an acoustic model, and, the output probability that the feature quantity of the sound signal subjected to retrieval is output from each phoneme state created from the acoustic model, wherein in the output probability obtaining step, when the retrieval string is converted into the phoneme string in the converting step, the output probability stored in association with each phoneme state contained in the phoneme string is obtained from the output probabilities stored in the output probability memory for each of the frames contained in the likelihood obtainment segment.
  16. 16
    The voice retrieval method according to claim 15, further comprising a replacing step of replacing each of the output probabilities obtained in the output probability obtaining step for each of the frames with a maximum output probability in the frame, an N1 number of frames previous to the frame, and an N2 number of frames subsequent to the frame, wherein: the symbols N1 and N2 are each a natural number including zero, but either the N1 or the N2 is not zero; and in the likelihood obtaining step, based on the output probability having undergone replacement in the replacing step, the likelihood of the likelihood obtainment segment specified in the segment specifying step is obtained.
  17. 17
    Independent claimA non-transitory recording medium having recorded therein a program that causes a computer of a voice retrieval apparatus including a memory recording a sound signal subjected to retrieval to execute: a converting process of converting a retrieval string into a phoneme string; a time length obtaining process of obtaining, from a database that stores continuous time length data on a phoneme, a continuous time length for each phoneme contained in the phoneme string converted in the converting process; a time length deriving process of deriving, based on the continuous time length obtained in the time length obtaining process, a plurality of time lengths corresponding to a plurality of utterance rates different from one another as candidate utterance time lengths of voices corresponding to the retrieval string; a segment specifying process of specifying, for each of the plurality of time lengths, a plurality of likelihood obtainment segments within the sound signal subjected to retrieval, each of the plurality of likelihood obtainment segments having the time length derived in the time length deriving process; a likelihood obtaining process of obtaining a likelihood showing a plausibility that the likelihood obtainment segment specified in the segment specifying process is a segment where voices corresponding to the retrieval string are uttered; an identifying process of identifying, based on the likelihood obtained in the likelihood obtaining process, an estimation segment where, within the sound signal subjected to retrieval, utterance of voices corresponding the retrieval string is estimated, the estimation segment being identified for each of the likelihood obtainment segments specified in the segment specifying process; and a displaying process of displaying, on a screen of an output device, the estimation segment identified in the identifying process, wherein: the program further causes the computer to execute a selecting process of selecting one of the plurality of time lengths based on the likelihood obtained in the likelihood obtaining process; and in the identifying process, the estimation segment is identified among the likelihood obtainment segments with the selected time length based on the likelihood obtained for the likelihood obtainment segment with the selected time length.
  18. 18
    Independent claimA non-transitory recording medium having recorded therein a program that causes a computer of a voice retrieval apparatus including a memory recording a sound signal subjected to retrieval to execute: a converting process of converting a retrieval string into a phoneme string; a time length obtaining process of obtaining, from a database that stores continuous time length data on a phoneme, a continuous time length for each phoneme contained in the phoneme string converted in the converting process; a time length deriving process of deriving, based on the continuous time length obtained in the time length obtaining process, a plurality of time lengths corresponding to a plurality of utterance rates different from one another as candidate utterance time lengths of voices corresponding to the retrieval string; a segment specifying process of specifying, for each of the plurality of time lengths, a plurality of likelihood obtainment segments within the sound signal subjected to retrieval, each of the plurality of likelihood obtainment segments having the time length derived in the time length deriving process; a likelihood obtaining process of obtaining a likelihood showing a plausibility that the likelihood obtainment segment specified in the segment specifying process is a segment where voices corresponding to the retrieval string are uttered; an identifying process of identifying, based on the likelihood obtained in the likelihood obtaining process, an estimation segment where, within the sound signal subjected to retrieval, utterance of voices corresponding the retrieval string is estimated, the estimation segment being identified for each of the likelihood obtainment segments specified in the segment specifying process; and a displaying process of displaying, on a screen of an output device, the estimation segment identified in the identifying process, wherein: the program further causes the computer to execute: a feature quantity calculating process of calculating a feature quantity of the sound signal subjected to retrieval in the likelihood obtainment segment specified in the segment specifying process for each frame that is a time window to compare the sound signal with an acoustic model; and an output probability obtaining process of obtaining, for each of the frames, an output probability that the feature quantity of the sound signal subjected to retrieval is output from each phoneme contained in the phoneme string, and in the likelihood obtaining process, respective values that are each a logarithm of the output probability obtained for each of the frames contained in the likelihood obtainment segment specified in the segment specifying process are added to obtain the likelihood of the likelihood obtainment segment.

Claim map

Independent claims stand on their own. The others add detail to the claim they name.

Claim 14 claims build on it
Claim 62 claims build on it
Claim 94 claims build on it
Claim 142 claims build on it
Claim 17No claims build on it
Claim 18No claims build on it

Description

Cross-reference to related application

This application claims the benefit of Japanese Patent Application No. 2014-259419, filed on Dec. 22, 2014, the entire disclosure of which is incorporated by reference herein.

Field

This application relates generally to a voice retrieval apparatus, a voice retrieval method, and a non-transitory recording medium.

Background

Due to widespread popularization of multimedia contents, such as voice and motion image, there is a demand for a highly precise multimedia retrieval technology. With respect to such a technology, a voice retrieval technology that identifies a portion where voices corresponding to a retrieval term (query) subjected to retrieval is uttered in a sound signal has been studied.

As for voice retrieval, a retrieval scheme with a sufficient performance has not been established yet in comparison with character string retrieval technologies based on image recognition. Hence, various technologies have been studied in order to realize a voice retrieval with a sufficient performance.

For example, Non-patent Literature 1 (Y. Zhang and J. Glass, “An inner-product lower-bound estimate for dynamic time warping”, in Proc., ICASSP, 2011, pp. 5660-5663) discloses a method of comparing sound signals with each other at a fast speed. This method enables a fast-speed identification of a portion corresponding to a query input by voice in a sound signal subjected to retrieval.

According to the technology disclosed by Non-patent Literature 1, when, however, the utterance rate of voice subjected to retrieval is different from the utterance rate of a person who has input a query, the retrieval precision decreases.

The present disclosure has been made in order to address the aforementioned technical problem, and it is an objective of the present disclosure to provide a voice retrieval apparatus, a voice retrieval method, and a non-transitory recording medium which are capable of highly precisely retrieving a retrieval term from a sound signal with a different utterance rate.

Summary

In order to accomplish the above objective, a voice retrieval apparatus according to an aspect of the present disclosure includes:

a processor; and

a memory that records a sound signal subjected to retrieval,

in which the processor executes following processes:

a converting process of converting a retrieval string into a phoneme string;

a time length obtaining process of obtaining, from a database that stores continuous time length data on a phoneme, a continuous time length for each phoneme contained in the phoneme string converted in the converting process;

a time length deriving process of deriving, based on the continuous time length obtained in the time length obtaining process, a plurality of time lengths corresponding to a plurality of utterance rates different from one another as candidate utterance time lengths of voices corresponding to the retrieval string;

a segment specifying process of specifying, for each of the plurality of time lengths, a plurality of likelihood obtainment segments within the sound signal subjected to retrieval, each of the plurality of likelihood obtainment segments having the time length derived in the time length deriving process;

a likelihood obtaining process of obtaining a likelihood showing a plausibility that the likelihood obtainment segment specified in the segment specifying process is a segment where voices corresponding to the retrieval string are uttered; and

an identifying process of identifying, based on the likelihood obtained in the likelihood obtaining process, an estimation segment where, within the sound signal subjected to retrieval, utterance of voices corresponding the retrieval string is estimated, the estimation segment being identified for each of the likelihood obtainment segments specified in the segment specifying process.

The present disclosure enables highly precise retrieval of a retrieve term from a sound signal that has a different utterance rate.

Brief description of the drawings

A more complete understanding of this application can be obtained when the following detailed description is considered in conjunction with the following drawings, in which:

FIG. 1 is a diagram illustrating a physical structure of a voice retrieval apparatus according to a first embodiment of the present disclosure;

FIG. 2 is a diagram illustrating a functional structure of the voice retrieval apparatus according to the first embodiment of the present disclosure;

FIG. 3 is a diagram to explain a phoneme state;

FIG. 4 is a diagram to explain how to derive an utterance time length corresponding to an utterance rate;

FIG. 5A is a waveform diagram of a sound signal subjected to retrieval, FIG. 5B is a diagram illustrating a frame set in the sound signal subjected to retrieval, and FIG. 5C is a diagram illustrating a likelihood obtainment segment specified in the sound signal subjected to retrieval;

FIG. 6 is a diagram illustrating an example way of performing Lower-Bounding on an output probability;

FIG. 7 is a diagram to explain a selection method of a selector for candidate segments;

FIG. 8 is a flowchart illustrating a flow of a voice retrieval process executed by the voice retrieval apparatus according to the first embodiment of the present disclosure;

FIG. 9 is a flowchart illustrating a flow of process of identifying a segment corresponding to a retrieval string;

FIG. 10 is a diagram illustrating a functional structure of a voice retrieval apparatus according to a second embodiment of the present disclosure;

FIG. 11A is a diagram to explain a method of a selector to select candidate segments after an obtained likelihood weight coefficient is multiplied, and FIG. 11B is a diagram shown an example weight coefficient;

FIG. 12 is a diagram to explain a selection method of the selector for candidate segments; and

FIG. 13A is a diagram illustrating the example maximum likelihood of a segment obtained by dividing a sound signal by the selector and shown for each utterance rate, and

FIG. 13B is a diagram illustrating an example comparison on a likelihood ranking corresponding to an utterance rate for each segment obtained by dividing the sound signal.

Detailed description

An explanation will be given of a voice retrieval apparatus according to embodiments of the present disclosure with reference to the figures. The same or equivalent part throughout the figures will be denoted by the same reference numeral. First Embodiment

A voice retrieval apparatus 100 according to a first embodiment includes, as illustrated in FIG. 1 , as a physical structure, a Read-Only Memory (ROM) 1 , a Random Access Memory (RAM) 2 , an external memory device 3 , an input device 4 , an output device 5 , a Central Processing Unit (CPU) 6 , and a bus 7 .

The ROM 1 stores a voice retrieval program. The RAM 2 is utilized as a work area for the CPU 6 .

The external memory device 3 includes, for example, a hard disk, and stores data on a sound signal subjected to analysis (hereinafter, referred to as a retrieval sound signal), a mono-phone model, a tri-phone model, and a phoneme time length to be explained later.

The input device 4 includes, for example, a keyboard and a voice recognition device. The input device 4 supplies, to the CPU 6 , text data that is a retrieval term input by a user. The output device 5 includes, for example, a screen like a liquid crystal display, a speaker, and the like. The output device 5 displays, on the screen, the text data output by the CPU 6 , and outputs, from the speaker, voice data.

The CPU 6 reads the voice retrieval program stored in the ROM 1 and to the RAM 2 , and executes the voice retrieval program, thereby accomplishing the functions to be explained later. The bus 7 connects the ROM 1 , the RAM 2 , the external memory device 3 , the input device 4 , the output device 5 , and the CPU 6 one another.

The voice retrieval apparatus 100 includes, as a functional structure, as illustrated in FIG. 2 , a sound signal memory 101 , a mono-phone model memory 102 , a tri-phone model memory 103 , a time length memory 104 , a retrieval string obtainer 111 , a converter 112 , a time length obtainer 113 , a time length deriver 114 , a segment specifier 115 , a feature quantity calculator 116 , an output probability obtainer 117 , a replacer 118 , a likelihood obtainer 119 , a repeater 120 , a selector 121 , a second converter 122 , a second output probability obtainer 123 , a second likelihood obtainer 124 , and an identifier 125 . The sound signal memory 101 , the mono-phone model memory 102 , the tri-phone model memory 103 , and the time length memory 104 are constructed in the memory area of the external memory device 3 .

The sound signal memory 101 stores the retrieval sound signal. Example retrieval sound signals are sound like news broadcasting, recorded sound of meeting, recorded sound of lecture meeting, and movie sound.

The mono-phone model memory 102 and the tri-phone model memory 103 store respective acoustic models. The acoustic model is modeled frequency characteristics of each phoneme that constructs a character string that is obtainable as a retrieval string. More specifically, the mono-phone model memory 102 stores an acoustic model (mono-phone model) on the basis of a mono-phone (a phoneme), and the tri-phone model memory 103 stores an acoustic model (tri-phone model) on the basis of tri-phones (3 phonemes).

The term phoneme is a unit of component that constructs voices uttered by a person who utters a term. For example, a term “category” contains 8 phonemes that are “k”, “a”, “t”, “e”, “g”, “o”, “r”, and “i”.

The mono-phone model is an acoustic model created for each phoneme, and is an acoustic model that does not depend on adjacent phonemes, that is, an acoustic model which has a fixed state transition relative to previous and subsequent phonemes. The tri-phone model is an acoustic model created for each 3 phonemes, and is an acoustic model that depends on adjacent phonemes, that is, an acoustic model that has the state transition taken into consideration relative to the previous and subsequent phonemes. The tri-phone model contains a larger amount of information than that of the mono-phone model. The voice retrieval apparatus 100 learns the mono-phone model and the tri-phone model through a general scheme, and stores the mono-phone model and the tri-phone model in the mono-phone model memory 102 and the tri-phone model memory 103 , respectively, and beforehand.

An example acoustic model applicable as the mono-phone model and the tri-phone model is, for example, a Hidden Markov Model (HMM) that is utilized in general voice recognition technologies. The HMM is a model to stochastically estimate, from a sound signal, phonemes that construct such a sound signal by a statistical scheme. Utilized as for the HMM is a standard pattern that has parameters which are a transition probability showing a fluctuation of a state in a time, and a probability (output probability) that a feature quantity which is input in each state is output.

The time length memory 104 stores, in a unit of each phoneme state, a continuous time length of each phoneme utilized in the acoustic model in a manner classified into groups, such as for each utterance rate, for each gender, for each age group, and for each utterance environment. The continuous time length of each phoneme is an average time length when each phoneme is uttered. The state of each phoneme is a unit obtained by subdividing each phoneme in a direction of time, and is equivalent to the minimum unit of the acoustic model. Each phoneme has a number of states defined beforehand. An explanation will be given of an example case in which the number of states defined for each phoneme is “3”. For example, a Japanese voice “a” contains a phoneme “a” that has three states as illustrated in FIG. 3 which are a first state “a1” including a start of utterance of this phoneme, a second state “a2” that is an intermediate state, and a third state “a3” including an end of utterance of this phoneme. That is, a phoneme is constructed by three states. When the number of all phonemes utilized in the acoustic model is Q, there are (3×Q) number of states. The voice retrieval apparatus 100 calculates, for each phoneme state, an average value of the continuous time length based on a large quantity of data on sound signals, and stores a calculation result in the time length memory 104 beforehand.

In this embodiment, the groups of phoneme continuous time length is classified into groups corresponding to five types of utterance rate which are “fast”, “slightly fast”, “normal”, “slightly slow”, and “slow”. The time length memory 104 classifies a large amount of sound data into groups of five types of utterance rate that are “fast”, “slightly fast”, “normal”, “slightly slow”, and “slow”, obtains an average phoneme continuous time length for each utterance rate group, and stores the continuous time length for each group corresponding to the utterance rate.

The retrieval string obtainer 111 obtains a retrieval string input by the user via the input device 4 . That is, the user gives, to the voice retrieval apparatus 100 , a character string (text) that is a retrieval term (query) to retrieve a portion of the retrieval sound signal where target voices are uttered.

The converter 112 arranges, in sequence, the phonemes of the mono-phone model that does not depend on the adjacent phonemes in accordance with the retrieval string obtained by the retrieval string obtainer 111 , and converts the retrieval string into a phoneme string. That is, the converter 112 arranges the phonemes (mono-phones) in sequence when each character is uttered in the same sequence as the characters contained in the retrieval string, thereby converting the retrieval string into a mono-phone phoneme string.

In this embodiment, an explanation will be given of an example case in which a Japanese term “category” is to be retrieved. When, as the retrieval string, a Japanese term “category” is input, such a term “category” contains 8 phonemes (mono-phones) that are “k”, “a”, “t”, “e”, “g”, “o”, “r”, and “i”, and thus the converter 112 creates a phoneme string that is “k, a, t, e, g, o, r, i”.

The time length obtainer 113 obtains, from the time length memory 104 , the average continuous time length for each phoneme state corresponding to five types of utterance rate. The time length deriver 114 obtains, from the time length obtainer 113 , the continuous time length for each phoneme state contained in the phoneme string output by the converter 112 . Next, based on the obtained continuous time length, a time length of voices (hereinafter, referred to as an utterance time length) corresponding to the retrieval string is derived.

More specifically, first, the time length deriver 114 obtains, from the group of phoneme continuous time lengths for “fast”, the continuous time lengths corresponding to eight phonemes that are “k, a, t, e, g, o, r, i”. More specifically, since each phoneme has three states, and data on the continuous time length is accumulated for each state, 24 pieces of data on continuous time length are obtained. Next, the obtained continuous time lengths are added together, thereby deriving an utterance time length for the “fast” utterance rate for the phoneme string “k, a, t, e, g, o, r, i”. Subsequently, 24 pieces of data on continuous time length are likewise obtained from the group of phoneme continuous time lengths for “slightly fast”, and the utterance time length for the “slightly fast” utterance rate is derived. Likewise, 24 pieces of data on continuous time length are obtained from each of the group of phoneme continuous time lengths for “normal”, the group of phoneme continuous time lengths for “slightly slow”, and the group of phoneme continuous time lengths for “slow”, and thus respective utterance time lengths are derived.

This will be explained in more detail with reference to FIG. 4 . The first row in FIG. 4 shows the 24 states of eight phonemes of a retrieval term which are “category”. The second row shows a value when the continuous time length corresponding to each phoneme state is obtained from the group of continuous time lengths for the “fast” utterance rate in the time length memory 104 . In addition, a value (515 ms) obtained by totalizing 24 continuous time lengths is the utterance time length for the “fast” utterance rate. The third row shows a value when the continuous time length corresponding to each phoneme state is obtained from the group of continuous time lengths for the “slightly fast” utterance rate in the time length memory 104 . In addition, a value (635 ms) obtained by totalizing 24 continuous time lengths is the utterance time length for the “slightly fast” utterance rate. Likewise, the utterance time length (755 ms) for the “normal” utterance rate, the utterance time length (875 ms) for the “slightly slow” utterance rate, and the utterance time length (995 ms) for the “slow” utterance rate are obtained.

That is, the voice retrieval apparatus 100 prepares, in the time length memory 104 beforehand, typical five types of continuous time length at the time of utterance for each phoneme state, and derives five types of utterance time length corresponding to the retrieval term.

Returning to FIG. 2 , the segment specifier 115 obtains the retrieval sound signal from the sound signal memory 101 , and specifies, from the header of the retrieval sound signal in sequence, the segment of the utterance time length derived by the time length deriver 114 as a likelihood obtainment segment. The term likelihood is an indicator that shows a similarity level between voices subjected to retrieval and the phoneme string corresponding to the retrieval string created from the acoustic model. In order to compare the phoneme string converted from the retrieval string with the sound signal, the segment specifier 115 takes out the sound signal portion within the specified likelihood obtainment segment, and divides the taken-out signal portion into frames corresponding to each phoneme state contained in the phoneme string. The segment specifier 115 associates, for each of five time lengths derived by the time length deriver 114 , each frame contained in the taken-out sound signal portion with each phoneme state contained in the phoneme string.

The term frame is a time window that has a time length corresponding to a phoneme state. More specifically, the frame that is set in the retrieval sound signal will be explained with reference to FIGS. 5A, 5B and 5C . FIG. 5A is a waveform diagram of retrieval sound signal with a time length T from the header to the last. The vertical axis represents an amplitude of waveform, while the horizontal axis represents a time t. FIG. 5B illustrates a frame that is set in the sound signal illustrated in FIG. 5A . The first line indicates a 0th frame string beginning from the header of the sound signal. Since the number of phonemes contained in the Japanese term “category” is eight and there are 24 phoneme states, the number of frames contained in the 0th frame string is 24. Since the phoneme continuous time length varies depending on the utterance rate, a frame length F also varies depending on the utterance rate. Hence, five frame strings corresponding to five types of utterance rate that are “fast”, “slightly fast”, “normal”, “slightly slow”, and “slow” are set in the 0th frame string beginning from the header of the sound signal. A first frame string in the second line is set so as to be shifted from the header of the sound signal by a predetermined shift length S. The first frame string also has 24 frames, and five frame strings are set in accordance with the types of utterance rate. Subsequently, the header position of the frame string is likewise shifted by the shift length S, and setting of five frame strings is made up to a (P−1)th frame string.

The shift length S is a length to define the precision of a retrieval position to retrieve at which position in the sound signal the retrieval term is present. The shift length S is set to a fixed and shorter value than the shortest frame length. In this embodiment, since shortest length of a phoneme state shown in FIG. 4 is 15 ms, the shift length S is set to be 5 ms shorter than that length.

FIG. 5C illustrates a likelihood obtainment segment specified by the segment specifier 115 in the retrieval sound signal. First of all, the segment specifier 115 specifies the segment of the 0th frame string containing the 24 frames and beginning from the header of the sound signal as a 0th likelihood obtainment segment with a time length L. Since there are five 0th frame strings corresponding to the types of utterance rate, five 0th likelihood obtainment segments are specified in accordance with the types of utterance rate. Next, the segment of the first frame beginning from the position shifted by the shift length S from the header of the sound signal is specified as a first likelihood obtainment segment. Five first likelihood obtainment segments are likewise specified. Likewise, likelihood obtainment segment up to the (P−1)th likelihood obtainment segment corresponding to the (P−1)th frame string is specified five by five in sequence.

Returning to FIG. 2 , the feature quantity calculator 116 calculates, for each frame, the feature quantity of the retrieval sound signal within the likelihood obtainment segment specified by the segment specifier 115 . The feature quantity is obtainable by combining a frequency-axis-system feature parameter obtained by converting sound data on the frequency axis with a power-system parameter obtained by calculating a square sum of the energy of sound data and a logarithm thereof.

For example, as is conventionally well-known, the feature quantity comprises a 38-dimensional vector quantity with a total of 38 components: 12 components of the frequency-axis-system feature parameter (12 dimensions) and 1 component of the power-system feature parameter (1 dimension); a difference between each component of the present window and the previous time window, that is, 12 components of the Δ frequency-axis-system feature parameter (12 dimensions) and 1 component of the Δ power-system feature parameter (1 dimension); and a difference between a difference of each component of the present time window and the previous time window, that is, 12 components of the ΔΔ frequency-axis-coordinate feature parameter.

Returning to FIG. 2 , the output probability obtainer 117 obtains, for each frame, a probability (output probability) that the feature quantity is output from each phoneme contained in the phoneme string based on the feature quantity calculated by the feature quantity calculator 116 . More specifically, the output probability obtainer 117 obtains, from the mono-phone model memory 102 , the mono-phone model, and compares the feature quantity in each frame calculated by the feature quantity calculator 116 with the mono-phone model in the corresponding state to this frame in the phoneme states contained in the phoneme string. Next, the probability that the feature quantity in each frame is output from the corresponding state is calculated.

The output probability obtainer 117 calculates, for each of five likelihood obtainment segments corresponding to the utterance rate and specified by the segment specifier 115 , the output probability for each of 24 frames contained in each likelihood obtainment segment.

The replacer 118 replaces each output probability obtained by the output probability obtainer 117 with the maximum output probability in the adjacent several previous and subsequent frames. This replacement process is called Lower-Bounding. This process is also performed for each of five likelihood obtainment segments.

More specifically, with reference to FIG. 6 , the Lower-Bounding will be explained. In FIG. 6 , a continuous line indicates an output probability obtained for each frame. The vertical axis indicates the height of the output probability so as to increase toward the bottom, and the horizontal axis indicates a time t. The replacer 118 replaces the output probability of each frame with the maximum output probability in this frame, N1 number of previous frames, and N2 number of subsequent frames. N1 and N2 are both natural numbers including zero, but either N1 or N2 is not zero. An explanation will be given of a case in which N1=N2=2. The output probability of the first frame in the frame string is replaced with the maximum output probability in the first frame, the subsequent second and third frames since there is no frame previous to the first frame. The output probability of the second frame is replaced with the maximum output probability in the previous first frame, the second frame, and subsequent third and fourth frames. The output probability of the third frame is replaced with the maximum output probability in the previous first and second frames, the third frame, and the subsequent fourth and fifth frames. In this way, the replacement process is performed up to the 24th frame. Upon replacement, the output probability indicated by the continuous line is converted to an output probability that has a small change in value along a time direction like an LB (Lower-Bounding) output probability indicated by a dashed line.

By such Lower-Bounding, an error between the continuous time length of each phoneme stored in the time length memory 104 and the actual continuous time length of sound signal, and an error between the utterance time length of voices corresponding to the retrieval string derived by the time length deriver 114 and the actual utterance time length of sound signal are reduced within previous and subsequent several frames.

The likelihood obtainer 119 obtains a likelihood that shows a plausibility of the likelihood obtainment segment specified by the segment specifier 115 where voices corresponding to the retrieval string are uttered based on the output probability having undergone the replacement process by the replacer 118 . More specifically, the likelihood obtainer 119 adds the values obtained by getting a logarithm of each output probability having undergone the replacement process across all frames from the header of the likelihood obtainment segment to the last, in this example, the 24 frames to obtain the likelihood of this likelihood obtainment segment. That is, the more the likelihood obtainment segment contains frames with a high output probability, the larger the likelihood to be obtained by the likelihood obtainer 119 becomes. This process is also performed for each of five likelihood obtainment segments corresponding to the types of utterance rate.

Note that this is a process of multiplying the output probabilities of the respective frames, and thus the output probabilities may be directly multiplied without getting a logarithm, or an approximation formula may be applied instead of the logarithm.

The repeater 120 controls the respective components so as to change the specified segment in the sound signal portion in the likelihood obtainment segment specified by the segment specifier 115 , and to cause the segment specifier 115 , the feature quantity calculator 116 , the output probability obtainer 117 , the replacer 118 , and the likelihood obtainer 119 to repeat respective processes.

More specifically, with reference to FIGS. 5B and 5C , under the control of the repeater 120 , the segment specifier 115 shifts the header position of the frame by the shift length S (=10 ms) to specify the first frame string, and newly specifies the segment of the first frame string as a first likelihood obtainment segment. Five first likelihood obtainment segments are specified corresponding to the types of utterance rate. Next, the repeater 120 causes the respective components to repeat the processes of the respective components from the feature quantity calculator 116 to the likelihood obtainer 119 in this newly specified first likelihood obtainment segment, thereby obtaining the likelihood of the first likelihood obtainment segment.

Likewise, the repeater 120 causes the segment specifier 115 to shift the specifying likelihood obtainment segment by the shift length S (=10 ms) from the second likelihood obtainment segment to the (P−1)th likelihood obtainment segment, thereby controlling the respective components to obtain the likelihood in each likelihood obtainment segment having undergone the shifting. Consequently, for each likelihood obtainment segment obtained by shifting the retrieval sound signal by the shift length S, the likelihood for the phoneme string “k, a, t, e, g, o, r, i” created based on the mono-phone model is obtained.

Note that a number P of the likelihood obtainment segments specifiable in the retrieval sound signal is defined as P=(T−L+S)/S where T is the time length of the sound signal, L is the time length of the likelihood obtainment segment, and S is the shift length. Since five likelihood obtainment segments are set for each type of utterance rate, the likelihood is obtained for each of a total of 5P number of likelihood obtainment segments.

The selector 121 selects, in the order of higher likelihood, x number of candidate estimate segments where the utterance of voice corresponding to the retrieval string is estimated in the likelihood obtainment segments specified by the segment specifier 115 based on the likelihood obtained by the likelihood obtainer 119 . That is, in order to reduce the calculation amount for a further precise likelihood obtainment based on the tri-phone model at a later stage, the selector 121 preliminary selects x number of segments that will be candidates of a final retrieval result from the 5P number of likelihood obtainment segments from which the respective likelihoods have been obtained, and excludes the remaining likelihood obtainment segments from the candidates.

At this time, since the likelihood obtainment segments specified by the segment specifier 115 have a large number of overlap portions, the segments where the likelihood is high often present in sequence in time series. Hence, when the selector 121 simply selects, in the order of higher likelihood, the candidate estimation segments among the likelihood obtainment segments, a possibility that the segments to be selected are concentrated at a portion of the retrieval sound signal increases.

In order to avoid this occasion, the selector 121 sets a predetermined selection time length, and for each selection time length, the likelihood obtainment segment with the maximum likelihood is selected one by one among the likelihood obtainment segments beginning within the predetermined selection time length. The predetermined selection time length is set to be shorter than the time length L of the likelihood obtainment segment like a time length corresponding to 1/m (for example, ½) of the time length L of the likelihood obtainment segment. When, for example, the utterance time length of the Japanese term “category” is assumed as being equal to or longer than 2 seconds (L≧2 seconds), the value m is set to be m=2, and the selection time length is set to be 1 second. The likelihood obtainment segment is selected one by one as a candidate for each selection time length (L/m), and the others are excluded from the candidates. Hence, the selector 121 is capable of sufficiently selecting the candidate estimation segments across the whole retrieval sound signal.

The selector 121 selects x number of likelihood obtainment segments with a high likelihood among the likelihood obtainment segments selected for each selection time length (L/m). This selection is performed for each of the likelihood obtainment segments corresponding to the five types of utterance rate. That is, for each of five types of utterance rate, x number (a total of 5×) of likelihood obtainment segments with a high likelihood are selected among the likelihood obtainment segments selected for each of the five types of utterance rate.

Next, in order to further limit the segments to be the candidates, the selector 121 compares the likelihood of the likelihood obtainment segments among the five types of utterance rate, selects only the x number of likelihood obtainment segments corresponding to the utterance rate that has a high likelihood as final candidates, and excludes the remaining likelihood obtainment segments from the candidates.

An explanation will be given of an example case in which the number x for selection is 10 (x=10) with reference to FIG. 7 . The utterance time length for the Japanese term “category” is assumed as being equal to or longer than 2 seconds, and the selection time length is set to 1 second. First, the selector 121 selects, one by one, the likelihood obtainment segment with the highest likelihood for each selection time length (1 second) among the P number of likelihood obtainment segments corresponding to the “fast” utterance rate. Next, 10 likelihood obtainment segments in the order of higher likelihood are selected from the likelihood obtainment segments selected for each 1 second, and the selected likelihood obtainment segments are stored in the likelihood field for the “fast” utterance rate in FIG. 7 . Next, the selector 121 totalizes the 10 likelihoods (0.93). Subsequently, the selector 121 selects, one by one, the likelihood obtainment segment with the highest likelihood for each selection time length (1 second) among the P number of likelihood obtainment segments corresponding to the “slightly fast” utterance rate. Next, 10 likelihood obtainment segments in the order of higher likelihood are selected from the likelihood obtainment segments selected for each 1 second, and the selected likelihood obtainment segments are stored in the likelihood field for the “slightly fast” utterance rate in FIG. 7 . Next, the selector 121 totalizes the 10 likelihoods (1.39). Likewise, the total value corresponding to the “normal” utterance rate is (2.12), the total value corresponding to the “slightly slow” utterance rate is (2.51), and the total value corresponding to the “slow” utterance rate is (1.33).

Next, the selector 121 compares the respective total values one another, and selects, as the final candidates, only the 10 likelihood obtainment segments corresponding to “slightly slow” and having the largest total value (2.51).

A selection result by the selector 121 is displayed to the exterior via the screen of the output device 5 . Next, the voice retrieval apparatus 100 executes a likelihood obtaining process with a higher precision on the x number of selected segments based on the tri-phone model and a Dynamic Programming (DP) matching technique. The DP matching is a scheme of selecting a state transition so as to maximize the likelihood in the analysis segment. As for the tri-phone model, the state transition relative to previous and subsequent phonemes needs to be taken into consideration. Hence, the state transition relative to previous and subsequent phonemes is determined in such a way that the likelihood of the likelihood obtainment segment is maximized by DP matching.

The second converter 122 arranges the phonemes of the tri-phone model that is the second acoustic model depending on adjacent phonemes in sequence in accordance with the retrieval string obtained by the retrieval string obtainer 111 , thereby converting the retrieval string into a tri-phone phoneme string that is a second phoneme string. When, for example, the Japanese term “category” is input as the retrieval string, the term “category” contains six tri-phones that are “k−a+t”, “a−t+e”, “t−e+g”, “e−g+o”, “g−o+r”, and “o−r+i”. Thus, the second converter 122 creates a tri-phone phoneme string containing such six tri-phones arranged in sequence. In addition, bi-phones “k+a” and “r−i” each containing two phonemes may be allocated to the beginning and the last. In this case, the bi-phone model may be stored in the external memory device 3 beforehand. Note that a phoneme located at the left side of the symbol “−” is located prior to a center phoneme, and a phoneme located at the right side of the symbol “+” is located subsequent to the center phoneme.

The second output probability obtainer 123 obtains, for each frame, the output probability that the feature quantity of retrieval sound signal in the 10 likelihood obtainment segments corresponding to the “slightly slow” utterance rate and selected as the candidate estimation segments by the selector 121 is output from each phoneme contained in the second phoneme string (tri-phone phoneme string) converted by the second converter 122 . More specifically, the second output probability obtainer 123 obtains, from the tri-phone model memory 103 , the tri-phone model, and compares the feature quantity in each frame calculated by the feature quantity calculator 116 with each tri-phone model contained in the tri-phone phoneme string. Next, the output probability that the feature quantity is output from each tri-phone in each frame is calculated.

The second likelihood obtainer 124 obtains, for each of the candidate segments limited by the selector 121 into the x number (10 segments), a second likelihood showing the plausibility that the selected 10 likelihood obtainment segments for the “slightly slow” utterance rate as the candidate estimation segment by the selector 121 are the segments where voices corresponding to the retrieval string are uttered. The second likelihood is obtained based on the tri-phone phoneme string that is the second phoneme string. Accordingly, the second likelihood is an indicator with a higher precision than the likelihood obtained by the likelihood obtainer 119 based on the mono-phone phoneme string.

The second likelihood obtainer 124 retrieves, for each frame contained in the second likelihood obtainment segment limited by the selector 121 , an association between the feature quantity of the sound signal and each tri-phone model contained in the tri-phone phoneme string by DP matching based on the output probability obtained by the second output probability obtainer 123 . Next, by adding values obtained by taking a logarithm of the output probability obtained for each frame in the likelihood obtainment segment selected by the selector 121 , the second likelihood in this segment is obtained.

The identifier 125 identifies, among the 10 candidate segments for the “slightly slow” utterance rate selected by the selector 121 , the estimation segment where the utterance of voices corresponding to the retrieval string is estimated in the retrieval sound signal based on the second likelihood obtained by the second likelihood obtainer 124 . For example, the identifier 125 identifies, as the estimation segments, a predetermined number of segments in the order of larger second likelihood obtained by the second likelihood obtainer 124 . Alternatively, the identifier 125 identifies, as the estimation segment, the segment that has the likelihood equal to or larger than a predetermined value. The positional information on the segment identified by the identifier 125 is displayed to the exterior via the screen of the output device 5 as a final retrieval result.

An explanation will be given of a voice retrieval process executed by the voice retrieval apparatus 100 employing the above physical structure and functional structure with reference to the flowchart of FIG. 8 .

The description continues in the full USPTO document.

Timeline & family

Timeline From USPTO dates

2016201720182019202020212022202320242025Application filedNov 30, 2015Application publishedJune 23, 2016Patent grantedSep 19, 20173.5-year fee paidMarch 19, 20217.5-year fee not paidMarch 19, 2025Patent expiredSep 19, 2025

Maintenance fees

Fees are due 3.5, 7.5 and 11.5 years after grant. This patent expired on September 19, 2025, so the fee marked "not paid" was the one that went unpaid.

3.5-year feeDue March 19, 2021Paid
7.5-year feeDue March 19, 2025Not paid
11.5-year feeDue March 19, 2029Never came due

US family 2 documents, by filing date

Published applicationUS 2016/0180839 A1

VOICE RETRIEVAL APPARATUS, VOICE RETRIEVAL METHOD, AND NON-TRANSITORY RECORDING MEDIUM

Filed Nov 2015 · published Jun 2016
Published application
This documentUS 9,767,790 B2

Voice retrieval apparatus, voice retrieval method, and non-transitory recording medium

Filed Nov 2015 · granted Sep 2017
Lapsed, fee not paid

Earlier publications, parents and continuations. None of them can still be enforced, or this patent would not be listed.

US patents it cites 2

Prior art cited by the examiner or applicant. Useful when you check your own idea for novelty.

Sources & verification

Verification

  • The USPTO Official Gazette of November 18, 2025 lists it as expired on September 19, 2025 for an unpaid maintenance fee.
  • It isn't on any reinstatement notice published since.
  • Its 1 US relative has also lapsed, expired or never issued.
  • Rechecked against USPTO records every day.
  • We check US rights only. Check foreign counterparts before selling abroad.

Confirm it yourself

  1. Open the file history on Patent Center.
  2. The status should read "Patent Expired Due to NonPayment of Maintenance Fees Under 37 CFR 1.362".
  3. Check the documents for any later petition to revive or reinstate.

Everything on this page comes from the documents linked above.

More in AI & Machine Learning

All AI & Machine Learning
Drawing from US 9,767,787 B2Lapsed, fee not paid5 drawings
AI & Machine Learning · US 9,767,787 B2

Artificial utterances for speaker verification

A method for speaker verification is disclosed.

Filed2014
LapsedSep 2025
OwnerInternational Business Machines Corporation
Drawing from US 9,767,791 B2Lapsed, fee not paid8 drawings
AI & Machine Learning · US 9,767,791 B2

Method and apparatus for exemplary segment classification

Method and apparatus for segmenting speech by detecting the pauses between the words and/or phrases, and to determine whether a particular time interval contains speech or non-speech, such as a pause.

Filed2013
LapsedSep 2025
OwnerSPEECH MORPHING SYSTEMS, INC.