Lapsed, fee not paid11 drawingsAmplification of dynamic checks through concurrency fuzzing
The subject disclosure relates to effective dynamic monitoring of an application executing in a computing system by increasing concurrency coverage.
US 8,533,698 B2 · Assignee: Microsoft Corporation · Inventors: Zhu; Weirong et al.
Sheet 1 of 4 from the published document. All sheets in the USPTO PDF
The present invention extends to methods, systems, and computer program products for optimizing execution of kernels. Embodiments of the invention include an optimization framework for optimizing runtime execution of kernels. During compilation, information about the execution properties of a kernel are identified and stored alongside the executable code for the kernel. At runtime, calling contexts access the information. The calling contexts interpret the information and optimize kernel execution based on the interpretation.
Background and Relevant Art Computer systems and related technology affect many aspects of society. Indeed, the computer system's ability to process information has transformed the way we live and work. Computer systems now commonly perform a host of tasks (e.g., word processing, scheduling, accounting, etc.) that prior to the advent of the computer system were performed manually. More recently, computer systems have been coupled to one another and to other electronic devices to form both wired and wireless computer networks over which the computer systems and other electronic devices can transfer electronic data. Accordingly, the performance of many computing tasks are distributed across a number of different computer systems and/or a number of different computing environments. In some environments, execution of a program is split between multiple processors within the same computer sys
1 of 4 drawing sheets so far from the published document, cropped to the drawing. Every sheet is in the USPTO PDF.
What the patent claimed, word for word. All of it is now free to use.
Not Applicable.
Background and Relevant Art
Computer systems and related technology affect many aspects of society. Indeed, the computer system's ability to process information has transformed the way we live and work. Computer systems now commonly perform a host of tasks (e.g., word processing, scheduling, accounting, etc.) that prior to the advent of the computer system were performed manually. More recently, computer systems have been coupled to one another and to other electronic devices to form both wired and wireless computer networks over which the computer systems and other electronic devices can transfer electronic data. Accordingly, the performance of many computing tasks are distributed across a number of different computer systems and/or a number of different computing environments.
In some environments, execution of a program is split between multiple processors within the same computer system. For example, some computer systems have multiple Central Processing Units ("CPUs"). During program execution, two or more of the CPUs can execute part of a program. Further, many computer systems also include other types of processors, some with relatively significant processing capabilities. For example, many computer systems include one or more Graphics Processing Units ("GPUs"). Some combinations of compilers are capable of compiling source code into executable code (e.g., C++ to Intermediate Representation ("IR") to High Level Share Language ("HLSL") bytecode) that executes in part on a CPU and executes in part on a GPU. Some types of instructions may even be better suited for execution on a GPU. Thus, source code can be specifically developed for mixed execution on one or more CPUs and one or more GPUs.
In the domain of technical computing, it is typical that computational intensive kernels are accelerated by special hardware or networks. Typically the developer will demarcate boundaries of such a computationally intensive kernel (hereinafter referred to simply as a "kernel"). The boundaries indicate to a compiler when the code for the kernel is to be compiled in special ways such as, for example, to a different instruction set (that of the accelerator) or to set up a call-return sequence to and from the GPU
In the construction of compilers that address compilation for mixed execution (whether it be for a GPU, an accelerator, etc), it is useful to have as part of the compilation process, information flow from the calling context into the called context (e.g., the kernel) (top-down information flow). It is also useful to have as part of the compilation process, information flow from the called context (e.g., the kernel) into calling context (bottom-up information flow). Bi-directional flow is typical in a compilation approach called "whole program optimization. However, whole program optimization is not always practical or reliable. Whole program optimization is not always practical because sometimes a compiler is not configured to flow information in one or more of the top-down or bottom-up directions. Whole program optimization is not always reliable since the flow of information can depend on heuristics.
The present invention extends to methods, systems, and computer program products for optimizing execution of kernels. In some embodiments, lower level code is generated so that kernel execution can be optimized at runtime. Program source code is accessed. The program source code includes an element that identifies part of the program source code as a kernel that is to be executed on a co-processor (e.g., a Graphical Processing Unit ("GPU") or other accelerator). The program source code also declares properties of the kernel.
The program source code is compiled into lower level code. The code element within the program source code is detected. In response to detecting the code element, the program source code is compiled into proxy code and into separate stub code. The proxy code is for execution in a context (e.g., on a central processing unit) that can invoke the stub code and the stub code is for execution on a co-processor.
Compilation of the source code includes generating the proxy code. The proxy code is configured to invoke the stub code in accordance with the declared properties of the kernel. The proxy code includes a descriptor for referencing any runtime optimization objects stored during compilation.
An intermediate representation of the stub code is analyzed to derive usage information about the declared properties of the kernel. The stub code is generated in accordance with the derived usage information. The derived usage information is stored in one or more runtime optimization objects alongside the stub code. The descriptor is linked to the one or more runtime optimization objects stored alongside the stub code to provide the proxy code with access to the derived usage information for making kernel optimization decisions at runtime.
In other embodiments, kernel execution is optimized at runtime. An execution command is received to execute lower level code. The lower level code includes proxy code for execution on a central processing unit, stub code for execution on co-processor, and one or more runtime optimization objects stored alongside the stub code. The one or more runtime optimization objects store derived usage information about the usage of kernel properties declared in program source code. The derived usage information was derived through analysis during generation of the stub code from program source code. The proxy code is configured to invoke the stub code in accordance with the declared kernel properties. The proxy code includes a descriptor linked to the one or more runtime optimization objects stored alongside the stub code.
In response to the execution command, the proxy code is executed on a central processing unit to invoke a call proxy. Execution of the proxy code includes using the descriptor to consult the derived usage information stored in the one or more runtime optimization objects. Execution of the proxy code includes making an optimization decision optimizing execution of the kernel based on the derived usage information. The optimization decision includes optimizing one or more of: invoking the stub code and passing data to the stub code. Execution of the proxy code includes invoking the stub code on a co-processor (e.g., a Graphical Processing Unit ("GPU") or other accelerator). Execution of the proxy code includes passing data to the stub code to dispatch the kernel on the co-processor.
The stub code is executed on the co-processor to invoke a call stub. Execution of the stub code includes receiving the data passed from the call proxy. Execution of the call stub includes dispatching the kernel on the co-processor in accordance with the formatted data.
This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
Additional features and advantages of the invention will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of the invention. The features and advantages of the invention may be realized and obtained by means of the instruments and combinations particularly pointed out in the appended claims. These and other features of the present invention will become more fully apparent from the following description and appended claims, or may be learned by the practice of the invention as set forth hereinafter.
In order to describe the manner in which the above-recited and other advantages and features of the invention can be obtained, a more particular description of the invention briefly described above will be rendered by reference to specific embodiments thereof which are illustrated in the appended drawings. Understanding that these drawings depict only typical embodiments of the invention and are not therefore to be considered to be limiting of its scope, the invention will be described and explained with additional specificity and detail through the use of the accompanying drawings in which:
FIG. 1 illustrates an example computer architecture that facilitates generating lower level code so that kernel execution can be optimized at runtime.
FIG. 2 illustrates an example computer architecture that facilitates optimizing kernel execution at runtime.
FIG. 3 illustrates a flow chart of an example method for optimizing lower level code during compilation of program source code used to generate the lower level code.
FIG. 4 illustrates a flow chart of an example method for optimizing execution of a kernel on a co-processor.
The present invention extends to methods, systems, and computer program products for optimizing execution of kernels. In some embodiments, lower level code is generated so that kernel execution can be optimized at runtime. Program source code is accessed. The program source code includes an element that identifies part of the program source code as a kernel that is to be executed on a co-processor (e.g., a Graphical Processing Unit ("GPU") or other accelerator). The program source code also declares properties of the kernel.
The program source code is compiled into lower level code. The code element within the program source code is detected. In response to detecting the code element, the program source code is compiled into proxy code and into separate stub code. The proxy code is for execution in a context (e.g., on a central processing unit) that can invoke the stub code and the stub code is for execution on a co-processor.
Compilation of the source code includes generating the proxy code. The proxy code is configured to invoke the stub code in accordance with the declared properties of the kernel. The proxy code includes a descriptor for referencing any runtime optimization objects stored during compilation.
An intermediate representation of the stub code is analyzed to derive usage information about the declared properties of the kernel. The stub code is generated in accordance with the derived usage information. The derived usage information is stored in one or more runtime optimization objects alongside the stub code. The descriptor is linked to the one or more runtime optimization objects stored alongside the stub code to provide the proxy code with access to the derived usage information for making kernel optimization decisions at runtime.
In other embodiments, kernel execution is optimized at runtime. An execution command is received to execute lower level code. The lower level code includes proxy code for execution on a central processing unit, stub code for execution on co-processor, and one or more runtime optimization objects stored alongside the stub code. The one or more runtime optimization objects store derived usage information about the usage of kernel properties declared in program source code. The derived usage information was derived through analysis during generation of the stub code from program source code. The proxy code is configured to invoke the stub code in accordance with the declared kernel properties. The proxy code includes a descriptor linked to the one or more runtime optimization objects stored alongside the stub code.
In response to the execution command, the proxy code is executed on a central processing unit to invoke a call proxy. Execution of the proxy code includes using the descriptor to consult the derived usage information stored in the one or more runtime optimization objects. Execution of the proxy code includes making an optimization decision optimizing execution of the kernel based on the derived usage information. The optimization decision includes optimizing one or more of: invoking the stub code and passing data to the stub code. Execution of the proxy code includes invoking the stub code on a co-processor (e.g., a Graphical Processing Unit ("GPU") or other accelerator). Execution of the proxy code includes passing data to the stub code to dispatch the kernel on the co-processor.
The stub code is executed on the co-processor to invoke a call stub. Execution of the stub code includes receiving the data passed from the call proxy. Execution of the call stub includes dispatching the kernel on the co-processor in accordance with the formatted data.
Embodiments of the present invention may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present invention also include physical and other computer-readable media for carrying or storing computer-executable instructions and/or data structures. Such computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computer-executable instructions are computer storage media (devices). Computer-readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, embodiments of the invention can comprise at least two distinctly different kinds of computer-readable media: computer storage media (devices) and transmission media.
Computer storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives ("SSDs") (e.g., based on RAM), Flash memory, phase-change memory ("PCM"), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.
A "network" is defined as one or more data links that enable the transport of electronic data between computer systems and/or modules and/or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and/or data links which can be used to carry or desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.
Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to computer storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a "NIC"), and then eventually transferred to computer system RAM and/or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that computer storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.
Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and/or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
Those skilled in the art will appreciate that the invention may be practiced in network computing environments with many types of computer system configurations, including combinations having one or more of: personal computers, desktop computers, laptop computers, message processors, hand-held devices, multi-processor systems (including systems with a one or more Central Processing Units ("CPUs") and one or more co-processors, for example, Graphical Processing Units ("GPUs") or accelerators), microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, pagers, routers, switches, and the like. The invention may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.
Embodiments of the invention include an optimization framework for optimizing runtime execution of kernels. Information about the properties of a kernel is identified during compilation and stored alongside the executable code for the kernel. The information is made available to calling contexts at runtime. At runtime, calling contexts interpret the information and optimize kernel execution based on the interpretation.
There are a variety of applications for the optimization framework. One application includes, identifying at compile time and utilizing at runtime, memory immutability information about the immutability of memory resources consumed by a kernel. At runtime, memory immutability information can be used to evaluate the potential to run kernels concurrently and the potential to cache the memory resource contents without staleness.
Another application includes, identifying at compile time and utilizing at runtime, parameter reduction information for parameters passed to a kernel. A calling convention may require specified parameters to be passed to a kernel. However, the kernel may not actually consume all of the specified data. Parameter reduction information can identify portions of the specified data that are not used. At runtime, parameter reduction information can be used to relieve a calling context from passing parameters that would go unused at the kernel.
An additional application includes, identifying at compile time and utilizing at runtime, expected performance characteristics of a kernel. A runtime can use expected performance characteristics to select an appropriate co-processor for a kernel. For example, there may be multiple co-processors each with a different performance profile. At runtime, expected performance characteristics of a kernel can be used to assist in matching the kernel to a co-processor that can optimally execute the kernel.
A further application includes, identifying at compile time and utilizing at runtime, whether a kernel is to use mathematical libraries and/or advanced floating point operations. At runtime, information about the use of mathematical libraries and/or advanced floating point operations can be used to determine if auxiliary data (e.g., stored mathematical tables) is to be supplied to a kernel. At runtime, information about the use of mathematical libraries and/or advanced floating point operations can also be used to identify co-processors that are capable and/or best suited for the kernel.
FIG. 1 illustrates an example computer architecture that facilitates generating lower level code so that kernel execution can be optimized at runtime. Referring to FIG. 1, computer architecture 100 includes compiler 101. Compiler 101 is connected to other components (or is part of) of a system bus and/or a network, such as, for example, a Local Area Network ("LAN"), a Wide Area Network ("WAN"), and even the Internet. Accordingly, compiler 101 as well as any other connected computer systems and their components, can create message related data and exchange message related data (e.g., Internet Protocol ("IP") datagrams and other higher layer protocols that utilize IP datagrams, such as, Transmission Control Protocol ("TCP"), Hypertext Transfer Protocol ("HTTP"), Simple Mail Transfer Protocol ("SMTP"), etc.) over the system bus and/or network.
Generally, complier 101 is configured to compile higher level code into lower level code.
As depicted, compiler 101 includes parser/semantic checker 102, code generator 103, and analyzer 107. Parser/semantic checker 102 is configured to receive statements and expressions of higher level code (e.g., written in C++, C++ extended for parallel environments, Visual Basic, etc.). Parser/semantic checker 102 can parse and semantically check statements and expressions of higher level code to identify different aspects/portions of the higher level code, including identifying different routines that are to be executed on different processors. Parser/semantic checker 102 can also identify compile time information (e.g., object formats, etc.) and/or runtime information (e.g., resource handles, object contents, a thread specification, for example, a number of threads, etc.) associated with different routines in the higher level code. Parser/semantic checker 102 can output an intermediate representation ("IR") of received source code.
In some embodiments, statements and expressions of higher level code include annotations and/or language extensions that are used to specify a section of program source code corresponding to a kernel. The kernel includes code that is to run on a co-processor. Parser/semantic checker 102 can identify transitions between "normal" execution on a "host" (i.e., execution on a CPU) and execution on a co-processor from these annotations and/or extensions. Parser/semantic checker 102 can represent kernel related code as a separate routine in IR.
Code generator 103 is configured to receive IR from parser/semantic checker 102. From the IR, code generator 103 can generate a plurality of different lower level instructions (e.g., DirectX.RTM./High Level Shader Language ("HLSL") bytecode) that correctly implement the statements and expressions of received higher level code.
Analysis module 107 is configured to access IR generated by parser/semantic checker 102. Analysis module 107 can analyze the IR to identify interesting aspects of lower level kernel code (e.g., identifying optimization information that can be used at runtime to optimize kernel execution). Identified interesting aspects can be packed in objects and the objects stored alongside the lower level code. Accordingly, optimization information for optimizing kernel execution can be identified at compile time and stored along with the executable code for the kernel.
In some embodiments, stub code (and kernel) are outlined and represented as a separate routine in compiler 101's IR of a source code parse tree. Then, the stub code is lowered into an intermediate representation, which is machine independent. At a later stage, machine code appropriate for an accelerator (e.g., a GPU) is generated out of the stub function intermediate representation. Analysis of IR can occur at this alter stage, to glean interesting aspects of the stub code and make them available to the proxy code.
FIG. 3 illustrates a flow chart of an example method 300 for generating lower level code so that kernel execution can be optimized at runtime. Method 300 will be described with respect to the components and data depicted in computer architecture 100.
Method 300 includes an act of accessing program source code, the program source code including an element that identifies part of the program source code as a kernel that is to be executed on a co-processor, the program source code also declaring properties of the kernel (act 301). For example, compiler 101 can access higher level code 111. Higher level code 111 includes code annotation 112 identifying code portion 116 as a kernel for execution on a co-processor (e.g., a GPU or other accelerator). Higher level code 111 also includes declared kernel properties 171. Declared properties 171 can be used declared properties that a user declares are to be used for the kernel. Declared properties 171 can indicate read/write properties for memory resources to be used by the kernel, can define objects to be passed to the kernel, can define that the kernel is to use mathematical libraries, can define that the kernel is to use floating operations, etc.
Higher level code can also indicate a separation between normal code and co-processor code using any variety of other code elements, such a, for example, special functions or special statements. For example, in C++ a special function can be called to indicate separation of normal code and co-processor code.
Method 300 includes an act of compiling the program source code into lower level code (act 302). For example, compiler 101 can compiler higher level code 111 into lower level code 121 (e.g, HLSL byte code).
An act 302 includes an act of detecting the code element within the program source code (act 303). For example, parser/semantic checker 102 can detect code annotation 112 within higher level code 111. Parser/semantic checker 102 can create intermediate representation 181 form higher level code 111. Parser/semantic checker 102 can split kernel related code into stub routine 172 and calling context code into proxy routine 173 in accordance with code annotation 112 (i.e., code annotation 112 demarks the boundary between kernel code and other code).
Act 302 can include in response to detecting the code element, compiling the program source code into proxy code and into separate stub code, the proxy code for execution in a context (e.g., on a central processing unit) that can invoke the stub code and the stub code for execution on a co-processor (act 304). For example, in response to detecting code annotation 112, compiler 101 can compile higher level code 111 into proxy code 122 for execution on a CPU and stub code 123 for execution on a co-processor.
Act 304 includes an act of generating the proxy code, the proxy code configured to invoke the stub code in accordance with the declared properties of the kernel, the proxy code including a descriptor for referencing any runtime optimization objects stored during compilation (act 305). For example, code generator 103 can generate proxy code 122. Proxy code 122 is configured to invoke stub code 123 in accordance with declared properties 171. Proxy code 122 includes descriptor 124 for referencing any runtime optimizations stored during compilation of higher level code 111.
Proxy code 122 also includes data copy code 131. Proxy code 122 is configured to use descriptor 124 to access kernel usage information, such as, for example, usage information 176. Proxy code 122 is also configured to access co-processor characteristics for any co-processors available for kernel execution.
Stub invocation helper library 128 is a fixed helper library. Stub invocation helper library 128 is configured to receive usage data (e.g., contained in optimization objects), device (e.g., co-processor) characteristics, and input data (buffers and parameters). Based on these inputs, stub invocation helper library 128 can (a) select the appropriate device (e.g., co-processor), (b) package the data for the selected device, and (c) dispatch a call to the device. Stub code 122 can collect these data (e.g., optimization objects and buffer and parameters data) and pass them in generic form to the stub invocation helper library 128.
Data copy code 131 is configured to execute as data copy code in a runtime. At runtime, the data copy code copies data to the co-processor were the data is available to for stub code to consume, based on kernel optimization decisions.
Act 304 includes an act of analyzing an intermediate representation of the stub code to derive usage information about the declared properties of the kernel (act 306). For example, analysis module 107 can analyze stub routine 172 to derive usage information 176 about declared properties 171. Usage information 176 can include an indication that declared properties 171 are used to a lesser extent than declared. For example, declared properties 171 can declare a specified memory resource for a kernel as read/write. However, analysis module 107 can determine that data is only read from the specified memory resource. Usage information 176 can reflect that the specified memory resource is only read from. Alternately, analysis module 107 can determine that the kernel completely overwrites memory resource and never reads from it. Usage information 176 can reflect that the specified memory resource is overwritten and not read from.
Declared properties 171 can also declare objects that are to be passed to a kernel. Analysis module 107 can determine that less than all of the declared objects are actually used by the kernel. Usage information 176 can reflect that less than all of the declared objects are used by the kernel.
Declared properties 171 can also declare that a kernel is to use mathematical libraries and/or floating pointer operations. Analysis module 107 can determine that mathematical libraries and/or floating pointer operations are actually used by the kernel. Usage information 176 can reflect that mathematical libraries and/or floating pointer operations are actually used by the kernel.
Analysis module 107 can also characterize performance aspects of a kernel, such as, for example, making an assertion whether the kernel is compute-bound or memory-bound. Usage information 176 can reflect characterized performance aspects of a kernel.
Act 304 includes an act of generating the stub code in accordance with the derived usage information (act 307). For example, code generator 103 can generate stub code 123 in accordance with usage information 176. Generating stub code in accordance with usage information 176 can include optimizing stub code 123 or altering stub code 123 for optimization based on usage information 176. For example, declared properties 171 may indicate that stub code 123 is to receive three variables. However, usage information 176 may indicate that one of the three variables goes unused. Accordingly, code generator 103 optimizes stub code 123 by removing code for receiving the unused variable.
Similarly, declared properties 171 may indicate a memory resource as read and write. However, usage information 176 may indicate that no values are ever written to the memory resource (e.g., after initialization to a specified value). Accordingly, code generator 103 can alter stub code 123 by defining the memory resource as read only. Defining the memory resource as read only can facilitate optimizations at runtime.
Likewise, declared properties 171 may indicate that stub code 123 uses a mathematical library. However, usage information 176 may indicate that the mathematical library is not actually used. Accordingly, code generator 103 optimizes stub code 123 to remove reference to auxiliary data used by the mathematical library.
As depicted, stub code 123 includes data receive code 127 and kernel dispatch code 132. Data receive code 127 is configured to execute as data receive code in a runtime. At runtime, the data receive code receives data from a call proxy. Kernel dispatch code 132 is configured to execute as kernel dispatch code in a runtime. At runtime, the kernel dispatch code dispatches a kernel to perform work on the kernel formatted data.
Act 304 includes an act of storing the derived usage information in one or more runtime optimization objects alongside the stub code (act 308). For example, analysis module 107 can store usage information 176 in runtime optimization objects 118 alongside stub code 123. Act 304 includes an act of linking the descriptor to the one or more runtime optimization objects stored alongside the stub code to provide the proxy code with access to the derived usage information for making kernel optimization decisions at runtime (act 309). For example, descriptor 124 can be linked to runtime optimization objects 118 to provide proxy code 123 with access to usage information 176 at runtime. Thus, although proxy code 122 was not compiled based on usage information 176, proxy code 122 can still access and utilize usage information 176 at runtime to facilitate appropriate interface with and optimizations of a kernel.
FIG. 2 illustrates an example computer architecture 200 that facilitates optimizing kernel execution at runtime. Referring to FIG. 2, computer architecture 200 includes CPU runtime 201, co-processor runtime 203, and co-processor runtime 204. CPU runtime 201, co-processor runtime 203, and co-processor runtime 204 can be connected one another and to other components over (or be part of) of a system bus and/or a network, such as, for example, a Local Area Network ("LAN"), a Wide Area Network ("WAN"), and even the Internet. Accordingly, CPU runtime 201, co-processor runtime 203, and co-processor runtime 204 as well as any other connected computer systems and their components, can create message related data and exchange message related data (e.g., Internet Protocol ("IP") datagrams and other higher layer protocols that utilize IP datagrams, such as, Transmission Control Protocol ("TCP"), Hypertext Transfer Protocol ("HTTP"), Simple Mail Transfer Protocol ("SMTP"), etc.) over the system bus and/or network.
FIG. 4 illustrates a flow chart of an example method 400 for optimizing kernel execution at runtime. Method 400 will be described with respect to the components and data depicted in computer architecture 200.
Method 400 includes an act of receiving an execution command to execute lower level code, the lower level code including proxy code for execution on a central processing unit, stub code for execution on co-processor, and one or more runtime optimization objects stored alongside the stub code (act 401). For example, computer architecture 200 can receive run command 262 to execute lower level code 121. During compilation of lower level code 121 proxy code 122 was generated for execution on a CPU and stub code was generated for execution on a co-processor (e.g., a GPU or other accelerator). Runtime optimization objects 118 are stored alongside stub code 123. As described, runtime optimization objects 118 store usage information 176 previously derived by analysis module 107. Also as described, proxy code 122 is configured to invoke stub code 123 in accordance with declared kernel properties 171.
Method 400 includes in response to the execution command, an act of executing the proxy code on one of one or more central processing units to invoke a call stub (act 402). For example, proxy code 122 can be executed within CPU runtime 201 to invoke call proxy 202. Data copy code 231 can be instantiated within call proxy 202.
Act 402 includes an act of using the descriptor to consult the derived usage information stored in the one or more runtime optimization objects (act 403). For example, call proxy 222 can use descriptor 124 to link 125 to usage information 176 stored in runtime optimization objects 118. Call proxy 202 can also access co-processor characteristics 276. Co-processor characteristics 276 indicate the execution characteristics, such as, for example, compute capabilities, memory interface capabilities, floating point capabilities, etc. of co-processor runtimes 203, 204, etc. Call proxy 202 can also call stub invocation helper library 128 to execute usage information analysis code 226 and stub code invocation code 228. Call proxy 202 can pass usage information 176 and co-processor characteristics 276 to usage information analysis code 226.
Act 402 includes an act of making one or more optimization decisions optimizing execution of the kernel based on the derived usage information, including optimizing one or more of: invoking the stub code and passing data to the stub code (act 404). For example, usage information analysis code 226 can make optimization decisions 272 for optimizing execution of a kernel based on usage information 176. Usage information analysis code 226 can also base optimization decisions on co-processor characteristics 276 or on a combination of usage information 176 and co-processor characteristics 276. Optimization decisions 272 can be passed to stub code invocation code 228 and data copy code 231.
An optimization decision can include sending less than all data that is otherwise indicated by declared properties 171. For example, it may be that declared properties 171 indicate that a kernel is to receive four objects as input. However, analysis module 107 can determine that two of the objects are not actually used by the kernel. Thus, usage information analysis code 226 can make an optimization decision to refrain from sending the two declared but unused objects to the kernel.
Similarly, it may be that declared properties 171 indicate that a kernel is to use mathematical libraries. Analysis module 107 can verify that the mathematical libraries are actually used by the kernel. Thus, usage information analysis code 226 can dictate that auxiliary information for the mathematical libraries be supplied to the kernel.
Further it may be that declared properties 171 indicate that data is both read from and written to a memory resource. However, analysis module 107 can determine that data is never actually written to the memory resource (e.g., the memory resource retains initialized data). Thus, usage information analysis code 226 can make an optimization decision to run multiple kernels concurrently for multiple kernels that just read the memory resource. Alternately, analysis module 107 can determine that a kernel completely overwrites the memory resources but the memory resource is never read form. Thus, usage information analysis code 226 can make an optimization decision not to initialize the memory resource prior to sending the memory resource to stub code.
Analysis module 107 can determine that a kernel has various performance aspects, such as, for example, that a kernel is compute-bound, that a kernel is memory-bound, that a kernel uses integer math, that a kernel uses floating point math, that a kernel has divergent control flow paths, that a kernel uses a memory access pattern (e.g., streaming or temporal locality), etc. Usage information analysis code 226 can make an optimization decision to invoke a kernel on an available co-processor that is appropriately suited to execute the kernel based on the kernel's determined performance aspects.
Thus, although proxy code 122 is not compiled to address and/or interface with optimizations that were compiled stub code 123 at compile time, at runtime call proxy 202 adjusts to address these stub code optimizations through reference to stored usage information 176.
The description continues in the full USPTO document.
About 6,086 words. The USPTO PDF has it with every drawing.
Fees are due 3.5, 7.5 and 11.5 years after grant. This patent expired on September 10, 2025, so the fee marked "not paid" was the one that went unpaid.
OPTIMIZING EXECUTION OF KERNELS
Filed Jun 2011 · published Dec 2012Optimizing execution of kernels
Filed Jun 2011 · granted Sep 2013Earlier publications, parents and continuations. None of them can still be enforced, or this patent would not be listed.
Prior art cited by the examiner or applicant. Useful when you check your own idea for novelty.
Everything on this page comes from the documents linked above.