ITU-R BS 1196-5-2015 Audio coding for digital broadcasting《数字传播音频编码》.pdf

资源描述

1、 Recommendation ITU-R BS.1196-5 (10/2015) Audio coding for digital broadcasting BS Series Broadcasting service (sound) ii Rec. ITU-R BS.1196-5 Foreword The role of the Radiocommunication Sector is to ensure the rational, equitable, efficient and economical use of the radio-frequency spectrum by all

2、radiocommunication services, including satellite services, and carry out studies without limit of frequency range on the basis of which Recommendations are adopted. The regulatory and policy functions of the Radiocommunication Sector are performed by World and Regional Radiocommunication Conferences

3、 and Radiocommunication Assemblies supported by Study Groups. Policy on Intellectual Property Right (IPR) ITU-R policy on IPR is described in the Common Patent Policy for ITU-T/ITU-R/ISO/IEC referenced in Annex 1 of Resolution ITU-R 1. Forms to be used for the submission of patent statements and lic

4、ensing declarations by patent holders are available from http:/www.itu.int/ITU-R/go/patents/en where the Guidelines for Implementation of the Common Patent Policy for ITU-T/ITU-R/ISO/IEC and the ITU-R patent information database can also be found. Series of ITU-R Recommendations (Also available onli

5、ne at http:/www.itu.int/publ/R-REC/en) Series Title BO Satellite delivery BR Recording for production, archival and play-out; film for television BS Broadcasting service (sound) BT Broadcasting service (television) F Fixed service M Mobile, radiodetermination, amateur and related satellite services

6、P Radiowave propagation RA Radio astronomy RS Remote sensing systems S Fixed-satellite service SA Space applications and meteorology SF Frequency sharing and coordination between fixed-satellite and fixed service systems SM Spectrum management SNG Satellite news gathering TF Time signals and frequen

7、cy standards emissions V Vocabulary and related subjects Note: This ITU-R Recommendation was approved in English under the procedure detailed in Resolution ITU-R 1. Electronic Publication Geneva, 2015 ITU 2015 All rights reserved. No part of this publication may be reproduced, by any means whatsoeve

8、r, without written permission of ITU. Rec. ITU-R BS.1196-5 1 RECOMMENDATION ITU-R BS.1196-5* Audio coding for digital broadcasting (Question ITU-R 19-1/6) (1995-2001-2010-2012-02/2015-10/2015) Scope This Recommendation specifies audio source coding systems applicable for digital sound and television

9、 broadcasting. It further specifies a system applicable for the backward compatible multichannel enhancement of digital sound and television broadcasting systems. Keywords Audio, Audio Coding, Broadcast, Digital, Broadcasting, Sound, Television, Codec The ITU Radiocommunication Assembly, considering

10、 a) that user requirements for audio coding systems for digital broadcasting are specified in Recommendation ITU-R BS.1548; b) that multi-channel sound system with and without accompanying picture is the subject of Recommendation ITU-R BS.775 and that a high-quality, multi-channel sound system using

11、 efficient bit rate reduction is essential in a digital broadcasting system; c) that the advanced sound system specified in Recommendation ITU-R BS.2051 consists of three dimensional channel configurations and uses either static or dynamic metadata to control audio objects; d) that subjective assess

12、ment of audio systems with small impairments, including multi-channel sound systems is the subject of Recommendation ITU-R BS.1116; e) that subjective assessment of audio systems of intermediate audio quality is subject of Recommendation ITU-R BS.1534 (MUSHRA); f) that low bit-rate coding for high q

13、uality audio has been tested by the ITU Radiocommunication Sector; g) that commonality in audio source coding methods among different services may provide increased system flexibility and lower receiver costs; h) that several broadcast services already use or have specified the use of audio codecs f

14、rom the families of MPEG-1, MPEG-2, MPEG-4, AC-3 and E-AC-3; i) that Recommendation ITU-R BS.1548 lists codecs that have been shown to meet the broadcasters requirements for contribution, distribution and emission; j) that those broadcasters which have not yet started services should be able to choo

15、se the system which is best suited to their application; k) that broadcasters may need to consider compatibility with legacy broadcasting systems and equipment when selecting a system; * This Recommendation should be brought to the attention of the International Standardization Organization (ISO) an

16、d the International Electrotechnical Commission (IEC). 2 Rec. ITU-R BS.1196-5 l) that when introducing a multi-channel sound system existing mono and stereo receivers should be considered; m) that a backward compatible multi-channel extension to an existing audio coding system can provide better bit

17、 rate efficiency than simulcast; n) that an audio coding system should preferably be able to encode both speech and music with equally high fidelity, recommends 1 that for new applications of digital sound or television broadcasting emission, where compatibility with legacy transmissions and equipme

18、nt is not required, one of the following low bit-rate audio coding systems should be employed: Extended HE AAC as specified in ISO/IEC 23003-3:2012; E-AC-3 as specified in ETSI TS 102 366 (2014-08); NOTE 1 Extended HE AAC is a more flexible superset of MPEG-4 HE AAC v2, HE AAC and AAC LC, and includ

19、es MPEG-D Unified Speech and Audio Coding (USAC). NOTE 2 E-AC-3 is a more flexible superset of AC-3. 2 that for applications of digital sound or television broadcasting emission, where compatibility with legacy transmissions and equipment is required, one of the following low bit-rate coding systems

20、 should be employed: MPEG-1 Layer II as specified in ISO/IEC 11172-3:1993; MPEG-2 Layer II half sample rate as specified in ISO/IEC 13818-3:1998; MPEG-2 AAC-LC or MPEG-2 AAC-LC with SBR as specified in ISO/IEC 13818-7:2006; MPEG-4 AAC-LC as specified in ISO/IEC 14496-3:2009; MPEG-4 HE AAC v2 as spec

21、ified in ISO/IEC 14496-3:2009; AC-3 as specified in ETSI TS 102 366 (2014-08); NOTE 3 ISO/IEC 11172-3 may sometimes be referred to as 13818-3 as this specification includes 11172-3 by reference. NOTE 4 The ITU-R Membership, as well as receiver and chipset manufacturers are encouraged to support Exte

22、nded HE AAC as specified in ISO/IEC 23003-3:2012. It includes all of the above mentioned AAC versions, thus guaranteeing compatibility with new future as well as legacy broadcast systems worldwide with the same single decoder implementation. 3 that for backward compatible multi-channel extension of

23、digital television and sound broadcasting systems, the multichannel audio extensions described in ISO/IEC 23003-1:2007 should be used; NOTE 5 Since the MPEG Surround technology described in ISO/IEC 23003-1:2007 is independent of the compression technology (core coder) used for transmission of the ba

24、ckward compatible signal, the described multi-channel enhancement tools can be used in combination with any of the coding systems recommended under recommends 1 and 2. 4 that for distribution and contribution links, ISO/IEC 11172-3 Layer II coding may be used at a bit rate of at least 180 kbit/s per

25、 audio signal (i.e. per mono signal, or per component of an independently coded stereo signal) excluding ancillary data; 5 that for commentary links, ISO/IEC 11172-3 Layer III coding may be used at a bit rate of at least 60 kbit/s excluding ancillary data for mono signals, and at least 120 kbit/s ex

26、cluding ancillary data for stereo signals, using joint stereo coding; 6 that for high quality applications the sampling frequency should be 48 kHz; Rec. ITU-R BS.1196-5 3 7 that the input signal to the low bit rate audio encoder should be emphasis-free and no emphasis should be applied by the encode

27、r; 8 that compliance with this Recommendation is voluntary. However, the Recommendation may contain certain mandatory provisions (to ensure e.g. interoperability or applicability) and compliance with the Recommendation is achieved when all of these mandatory provisions are met. The words “shall” or

28、some other obligatory language such as “must” and the negative equivalents are used to express requirements. The use of such words shall in no way be construed to imply partial or total compliance with this Recommendation, further recommends 1 that Recommendation ITU-R BS.1548 should be referred to

29、for information about coding system configurations that have been demonstrated to meet quality and other user requirements for contribution, distribution, and emission; 2 that further studies of the requirements for the advanced sound system specified in Recommendation ITU-R BS.2051 are needed and t

30、hat this Recommendation should be updated when these studies are completed. NOTE 1 Information about the codecs included in this Recommendation may be found in Appendices 1 to 5. Annex 1 (informative) MPEG-1 and MPEG-2, layer II and III audio 1 Encoding The encoder processes the digital audio signal

31、 and produces the compressed bit stream. The encoder algorithm is not standardized and may use various means for encoding, such as estimation of the auditory masking threshold, quantization, and scaling (following Note 1). However, the encoder output must be such that a decoder conforming to this Re

32、commendation will produce an audio signal suitable for the intended application. NOTE 1 An encoder complying with the description given in Annexes C and D to ISO/IEC 11172-3, 1993 will give a satisfactory minimum standard of performance. The following description is of a typical encoder, as shown in

33、 Fig. 1. Input audio samples are fed into the encoder. The time-to-frequency mapping creates a filtered and sub-sampled representation of the input audio stream. The mapped samples may be either sub-band samples (as in Layer I or II, see below) or transformed sub-band samples (as in Layer III). A ps

34、ycho-acoustic model, using a fast Fourier transform, operating in parallel with the time-to-frequency mapping of the audio signal creates a set of data to control the quantizing and coding. These data are different depending on the actual coder implementation. One possibility is to use an estimation

35、 of the masking threshold to control the quantizer. The scaling, quantizing and coding block creates a set of coded symbols from the mapped input samples. Again, the transfer function of this block can depend on the implementation of the encoding system. The block “frame packing” assembles the actua

36、l bit stream for the chosen layer from the output data of the other blocks (e.g. bit allocation data, scale factors, coded sub-band samples) and adds other information in the ancillary data field (e.g. error protection), if necessary. 4 Rec. ITU-R BS.1196-5 FIGURE 1 Block diagram of a typical encode

37、r BS .11 96 -01P C Ma udi o s i gna lT i m e - t o- f r e que nc ym a ppi ngS c a l i ngqua nt i z i nga nd c odi ngF r a m e pa c ki ngP s yc hoa c ous t i c m ode lI S O / I E C 1 1 172 - 3c ode d b i t s t r e a mI S O / I E C 1 1 172 - 3 e nc ode d A nc i l l a r y d a t a2 Layers Depending on t

38、he application, different layers of the coding system with increasing complexity and performance can be used. Layer I: This layer contains the basic mapping of the digital audio input into 32 sub-bands, fixed segmentation to format the data into blocks, a psycho-acoustic model to determine the adapt

39、ive bit allocation, and quantization using block companding and formatting. One Layer I frame represents 384 samples per channel. Layer II: This layer provides additional coding of bit allocation, scale factors, and samples. One Layer II frame represents 3 384 = 1 152 samples per channel. Layer III:

40、 This layer introduces increased frequency resolution based on a hybrid filter bank (a 32 sub-band filter bank with variable length modified discrete cosine transform). It adds a non-uniform quantizer, adaptive segmentation, and entropy coding of the quantized values. One Layer III frame represents

41、1 152 samples per channel. There are four different modes possible for any of the layers: single channel; dual channel (two independent audio signals coded within one bit stream, e.g. bilingual application); stereo (left and right signals of a stereo pair coded within one bit stream); joint stereo (

42、left and right signals of a stereo pair coded within one bit stream with the stereo irrelevancy and redundancy exploited). The joint stereo mode can be used to increase the audio quality at low bit rates and/or to reduce the bit rate for stereophonic signals. Rec. ITU-R BS.1196-5 5 3 Coded bit strea

43、m format An overview of the ISO/IEC 11172-3 bit stream is given in Fig. 2 for Layer II and Fig. 3 for Layer III. A coded bit stream consists of consecutive frames. Depending on the layer, a frame includes the following fields: FIGURE 2 ISO/IEC 11172-3 Layer II bit stream format BS .11 96 -02F ra m e

44、 1n F r a m e n F r a m e + 1nA nc i l l a r y d a t aM a i n a udi o i nf or m a t i onL a ye r I I :pa r t of t he bi t s t r e a m c ont a i ni ng s ync hr oni z a t i on a nd s t a t usi nf or m a t i onpa r t of t he bi t s t r e a m c ont a i ni ng bi t a l l oc a t i on a nd s c a l e f a c t

45、 ori nf or m a t i onpa r t of t he bi t s t r e a m c ont a i ni ng e nc ode d s ub- ba nd s a m pl e spa r t of t he bi t s t r e a m c ont a i ni ng us e r de f i na bl e da t aH e a de rS i de i nf or m a t i onH e a de r :S i de i nf or m a t i on:M a i n a udi o i nf or m a t i on:A nc i l l a

46、 r y d a t a :6 Rec. ITU-R BS.1196-5 FIGURE 3 ISO/IEC 11172-3 Layer III bit stream format BS .1196-03L e ngt h_1 + L e ngt h_SI + L e ngt h_2SI SI SIH e a de r L e ngt h_1M a i n a udi o i nf or m a t i ona nc i l l a r y d a t aL a ye r I I I :S i de i nf or m a t i on ( S I ) :H e a de r :P oi nt

47、e r :L e ngt h_1 :M a i n a udi o i nf or m a t i on:A nc i l l a r y d a t a :L e ngt h_2 :pa r t of t he bi t s t r e a m c ont a i ni ng he a de r , po i nt e r , l e ngt h_1 a ndl e ngt h_2 , s c a l e f a c t or i nf or m a t i on, e t c .;pa r t of t he bi t s t r e a m c ont a i ni ng s ync h

48、r oni z a t i on a nd s t a t usi nf or m a t i on;poi nt i ng t o b e gi nni ng of m a i n a udi o i nf or m a t i on;l e ngt h o f f i r s t pa r t of m a i n a udi o i nf or m a t i on;l e ngt h o f s e c ond pa r t of m a i n a udi o i nf or m a t i on;pa r t of t he bi t s t r e a m c ont a i n

49、i ng e nc ode d a udi o;pa r t of t he bi t s t r e a m c ont a i ni ng us e r de f i na bl e da t a .P oi nt e r L e ngt h_24 Decoding The decoder accepts coded audio bit streams in the syntax defined in ISO/IEC 11172-3, decodes the data elements, and uses the information to produce digital audio output. The coded audio bit stream is fed into the decoder. The bit stream unpacking and decoding process optionally performs error detection if error-check is applied in the encoder. The bit stream is unpacked to recover the various pieces of inform

展开阅读全文