Unit-05/Lecture-01

Market-based management of Clouds

 

Market-based management of Clouds

·   Cloud computing already embodies the concept of providing IT assets as utilities. Firstly, it is important to understand what we intend by the term market which makes cloud computing different from market-oriented cloud computing.

·   The Oxford English Dictionary (OED) defines a market as a “place where a trade is conducted”. More precisely, market refers to a meeting or a gathering together of people for the purchase and sale of goods. A broader characterization defines the term market as the action of buying and selling, a commercial transaction, a purchase, or a bargain. Therefore, essentially the word market is the act of trading mostly performed in an environment—either physical or virtual that is specifically dedicated to such activity.

·   What differentiates market-oriented cloud computing (MOCC) from cloud computing is the presence of a virtual market place where IT services are traded and brokered dynamically. This is something that still has to be achieved and that will significantly evolve the way cloud computing services are eventually delivered to the consumer. More precisely, what is missing is the availability of a market where desired services are published and then automatically bid on by matching the requirements of customers and providers.

·   We can clearly characterize the relationship between cloud computing and MOCC as follows:

“Market Oriented Computing has the same characteristics as Cloud Computing; therefore it is a dynamically provisioned unified computing resource allowing you to manage software and data storage as on aggregate capacity resulting in “real-time” infrastructure across public and private infrastructures. Market Oriented Cloud Computing goes one step further by allowing spread into multiple public and hybrid environments dynamically composed by trading service.”

 

 

·   The realization of this vision is technically possible today but is not probable, given the lack of standards and overall immaturity of the market. Nonetheless, it is expected that in the near future, with the introduction of standards, concerns about security and trust will begin to disappear and enterprises will feel more comfortable leveraging a market-oriented model for integrating IT infrastructure and services from the cloud. Moreover, the presence of a demand-based marketplace represents an opportunity for enterprises to shape their infrastructure for dynamically reacting to workload spikes and for cutting maintenance costs. It also allows the possibility to temporarily lease some in-house capacity during low usage periods, thus giving a better return on investment. These developments will lead to the complete realization of market-oriented cloud computing.

 

A reference model for MOCC

Market-oriented cloud computing originated from the coordination of several components: service consumers, service providers, and other entities that make trading between these two groups possible. Market orientation not only influences the organization on the global scale of the cloud computing market. It also shapes the internal architecture of cloud computing providers that need to support a more flexible allocation of their resources, which is driven by additional parameters such as those defining the quality of service.

 

·   A global view of market-oriented cloud computing

A reference scenario that realizes MOCC at a global scale is given in Figure 5.1. It provides guidance on how MOCC can be implemented in practice.

 

·         Several components and entities contribute to the definition of a global market-oriented architecture. The fundamental component is the virtual marketplace—represented by the Cloud Exchange (CEx)—which acts as a market maker, bringing service producers and consumers together. The principal players in the virtual marketplace are the cloud coordinators and the cloud brokers. The cloud coordinators represent the cloud vendors and publish the services that vendors offer. The cloud brokers operate on behalf of the consumers and identify the subset of services that match customers’ requirements in terms of service profiles and quality of service. Brokers perform the same function as they would in the real world: They mediate between coordinators and consumers by acquiring services from the first and subleasing them to the latter. Brokers can accept requests from many users.

 

 At the same time, users can leverage different brokers. A similar relationship can be considered between coordinators and cloud computing services vendors. Coordinators take responsibility for publishing and advertising services on behalf of vendors and can gain benefits from reselling services to brokers. Every single participant has its own utility function that they all want to optimize rewards. Negotiations and trades are carried out in a secure and dependable environment and are mostly driven by SLAs, which each party has to fulfill. There might be different models for negotiation among entities, even though the auction model seems to be the more appropriate in the current scenario. The same consideration can be made for the pricing models: Prices can be fixed, but it is expected that they will most likely change according to market conditions.

 

·         Several components contribute to the realization of the Cloud Exchange and implement its features. In the reference model depicted in Figure 5.1, it is possible to identify three major components:

 

1.Directory

2.Auctioneer

3.Bank

 

 

 

 

            

Figure 5.1 Market-oriented cloud computing scenario

 

Ř  Directory. The market directory contains a listing of all the published services that are available in the cloud marketplace. The directory not only contains a simple mapping between service names and the corresponding vendor (or cloud coordinators) offering them. It also provides additional metadata that can help the brokers or the end users in filtering from among the services of interest those that can really meet the expected quality of service. Moreover, several indexing methods can be provided to optimize the discovery of services according to various criteria. This component is modified in its content by service providers and queried by service consumers.

 

Ř   Auctioneer. The auctioneer is in charge of keeping track of the running auctions in the marketplace and of verifying that the auctions for services are properly conducted and that malicious market players are prevented from performing illegal activities.

 

Ř   Bank. The bank is the component that takes care of the financial aspect of all the operations happening in the virtual marketplace. It also ensures that all the financial transactions are carried out in a secure and dependable environment. Consumers and providers may register with the bank and have one or multiple accounts that can be used to perform the transactions in the virtual marketplace.

This organization, as described, constitutes only a reference model that is used to guide system architects and designers in laying out the foundations of a Cloud Exchange system. In reality, the architecture of such a system is more complex and articulated since other elements have to be taken into account. For instance, since the cloud marketplace supports trading, which ultimately involves financial transactions between different parties, security becomes of fundamental importance. It is then important to put in place all the mechanisms that enable secure electronic transactions.

 

-----------------------------REFERENCE {book: M.K. Mastering of Cloud Computing, May 2013, Page No.  378-381}

                        

 

 

S.NO

RGPV QUESTIONS

Year

Marks

1

What is market oriented cloud computing? List the main component that implements a MOOC system.

 Dec 2014

7

 

 

 

 

 

 

 

Unit-05/Lecture-02

Market-based management of Clouds (….continued)

 

·         Market-oriented architecture for datacenters

Datacenters are the building blocks of the computing infrastructure that backs the services offered by a cloud computing vendor, no matter its specific category (IaaS, PaaS, or SaaS). In this section, we present these systems by taking into account the elements that are fundamental for realizing computing infrastructures that support MOCC. These criteria govern the logical organization of these systems—rather than their physical layout and hardware characteristics—and provide guidance for designing architectures that are market oriented. In other words, we describe reference architecture for MOCC datacenters.

 

Figure 5.2 provides an overall view of the components that can support a cloud computing provider in making available its services on a market-oriented basis. More specifically, the model applies to PaaS and IaaS providers that explicitly leverage virtualization technologies to serve customers’ needs.

 

There are four major components of the architecture:

Ř  Users and brokers: They originate the workload that is managed in the cloud data center. Users either require virtual machine instances to which to deploy their systems (IaaS scenario) or deploy applications in the virtual environment made available to them by the provider (PaaS scenario).These service requests are issued by service brokers that act on behalf of users and look for the best deal for them.

Ř   SLA resource allocator: The allocator represents the interface between the data center and the cloud service provider and the external world. Its main responsibility is ensuring that service requests are satisfied according to the SLA agreed to with the user. Several components coordinate allocator activities in order to realize this goal.

 

            

Figure 5.2 Reference architecture for a cloud datacenter

 

Ř  Service Request Examiner and Admission Control Module: This module operates in the front-end and filters user and broker requests in order to accept those that are feasible given the current status of the system and the workload that is already processing. Accepted requests are allocated and scheduled for execution. IaaS service providers allocate one or more virtual machine instances and make them available to users. PaaS providers identify a suitable collection of computing nodes to which to deploy the users’ applications.

Ř  Pricing Module: This module is responsible for charging users according to the SLA they signed. Different parameters can be considered in charging users; for instance, the most common case for IaaS providers is to charge according to the characteristics of the virtual machines requested in terms of memory, disk size, computing capacity, and the time they are used. It is very common to calculate the usage in time blocks of one hour, but several other pricing schemes exist. PaaS providers can charge users based on the number of requests served by their application or the usage of internal services made available by the development platform to the application while running.

Ř  Accounting Module: This module maintains the actual information on usage of resources and stores the billing information for each user. These data are made available to the Service Request Examiner and Admission Control module when assessing users’ requests. In addition, they constitute a rich source of information that can be mined to identify usage trends and improve the vendor’s service offering.

• Dispatcher. This component is responsible for the low-level operations that are required to realize admitted service requests. In an IaaS scenario, this module instructs the infrastructure to deploy as many virtual machines as are needed to satisfy a user’s request. In a PaaS scenario, this module activates and deploys the user’s application on a selected set of nodes; deployment can happen either within a virtual machine instance or within an appropriate sandboxed environment.

Ř  Resource Monitor: This component monitors the status of the computing resources, either physical or virtual. IaaS providers mostly focus on keeping track of the availability of VMs and their resource entitlements. PaaS providers monitor the status of the distributed middleware, enabling the elastic execution of applications and loading of each node.

Ř  Service Request Monitor: This component keeps track of the execution progress of service requests. The information collected through the Service Request Monitor is helpful for analyzing system performance and for providing quality feedback about the provider’s capability to satisfy requests. For instance, elements of interest are the number of requests satisfied versus the number of incoming requests, the average processing time of a request, or its time to execution. These data are important sources of information for tuning the system. The SLA allocator executes the main logic that governs the operations of a single datacenter or a collection of datacenters. Features such as failure management are most likely to be addressed by other software modules, which can either be a separate layer or can be integrated within the SLA resource allocator.

Ř  Virtual machines (VMs): Virtual machines constitute the basic building blocks of a cloud computing infrastructure, especially for IaaS providers. VMs represent the unit of deployment for addressing users’ requests. Infrastructure management software is in charge of keeping operational the computing infrastructure backing the provider’s commercial service offering. As we discussed, VMs play a fundamental role in providing an appropriate hosting environment for users’ applications and, at the same time, isolate application execution from the infrastructure, thus preventing applications from harming the hosting environment. Moreover, VMs are among the most important components influencing the QoS with which a user request is served.

VMs can be tuned in terms of their emulated hardware characteristics so that the amount of computing resource of the physical hardware allocated to a user can be finely controlled. PaaS providers do not directly expose VMs to the final user, but they may internally leverage virtualization technology in order to fully and securely utilize their own infrastructure. As previously discussed PaaS providers of ten leverage given middleware for executing user applications and might use different QoS parameters to charge application execution rather than the emulated hardware profile.

Ř  Physical machines: At the lowest level of the reference architecture resides the physical infrastructure that can comprise one or more data centers. This is the layer that provides the resources to meet service demands.

 

This architecture provides cloud services vendors with a reference model suitable to enabling their infrastructure for MOC. As mentioned, these observations mostly apply to PaaS and IaaS pro- viders, whereas SaaS vendors operate at a higher abstraction level. Still, it is possible to identify some of the elements of the SLA resource allocator, which will be modified to deal with the ser- vices offered by the provider. For instance, rather than linking user requests to virtual machine instances and platform nodes, the allocator will be mostly concerned with scheduling the execution of requests within the provider’s SaaS framework, and lower layers in the technology stack will be in charge of controlling the computing infrastructure. Accounting, pricing, and service request monitoring will still perform their roles.

 

Technologies and initiatives supporting MOCC

Existing cloud computing solutions have very limited support for market-oriented strategies to deliver services to customers. Most current solutions mainly focused on enabling cloud computing concern the delivery of infrastructure, distributed runtime environments, and services. Since cloud computing has been recently adopted, the consolidation of the technology constitutes the first step toward the full realization of its promise. Until now, a good deal of interest has been directed toward IaaS solutions, which represent a well-consolidated sector in the cloud computing market, with several different players and competitive offers. New PaaS solutions are gaining momentum, but it is harder for them to penetrate the market dominated by giants such as Google, Microsoft, and Force.com.

·   Framework for trading computing utilities

-          From an academic point of view, a considerable amount of research has been carried out in defining models that enable the trading of computing utilities, with a specific focus on the design of market-oriented schedulers for grid computing systems.

-           Computing grids aggregate a heterogeneous set of resources that are geographically distributed and might belong to different organizations. Such resources are often leased for long-term use by means of agreements among these organizations.

-           Within this context, market-oriented schedulers, who are aware of the price of a given computing resource and schedule user’s applications according to their budgets, have been investigated and implemented. The research in this area is of relevance to MOCC, since cloud computing leverages preexisting distributed computing technologies, including grid computing.

-          Garg and Buyya have provided a complete taxonomy and analysis of such schedulers, which is reported in Figure 5.3. A major classification categorizes these schedulers according to allocation decision, objective, market model, application model, and participant focus. Of particular interest is the classification according to the market model, which is the mechanism used for trading between users and providers.

 

Along this dimension, it is possible to classify the schedulers into the following categories:

-          Game theory: In market models that are based on game theory, participants interact in the form of an allocation game, with different payoffs as a result of specific actions that employ various strategies.

-          Proportional share: This market model originates from proportional share scheduling, which aims to allocate jobs fairly over a set of resources. This original concept has been contextualized within a market-oriented scenario in which the shares of the cluster are directly proportional to the user’s bid.

-          Commodity market: In this model the resource provider specifies the price of resources and charges users according to the amount of resources they consume. The provider’s determination of the price is the result of a decision process involving investment and management costs, current demand, and supply. Moreover, prices might be subject to vary over time.

-          Posted price: This model is similar to the commodity market, but the provider may make special offers and discounts to new clients. Furthermore, with respect to the commodity market, prices are fixed over time.

-          Contract-Net.: In market models based on the Contract-Net protocol, users advertise their demand and invite resource owners to submit bids. Resource owners check these advertisements with respect to their requirements. If the advertisement is favorable to them, the providers will respond with a bid. The user will then consolidate all the bids and compare them to select those most favorable to him. The providers are then informed about the outcome of their bids, which can be acceptance or rejection.

 

                

Figure 5.3 Market-oriented scheduler taxonomy

 

• Bargaining. In market models based on bargaining, the negotiation among resource consumers and providers is carried out until a mutual agreement is reached or it is stopped when either of the parties is no longer interested.

-          Auction: In market models based on auctions, the price of resources is unknown, and competitive bids regulated by a third party—the auctioneer—contribute to determining the final price of a resource. The bid that ultimately sets the price of a resource is the winning bid, and the corresponding user gains access to the resource.

The most popular and interesting market models for trading computing utilities are the commodity market, posted price, and auction models. Commodity market and posted price models, or variations/combinations of them, are driving the majority of cloud computing services offerings today. Auction-based models can instead potentially constitute the reference market-models for MOCC, since they are able to seamlessly support dynamic negotiations.

 

                                

 

 

 

 

 

-----------------------------REFERENCE {book: M.K. Mastering of Cloud Computing, May 2013, Page No.  381-387}

 

 

 

 

 

S.NO

RGPV QUESTIONS

Year

Marks

1

What are the main components that implement a MOOC based system.

Dec 2014, June 2015

7

 

 

 

 

 

 

 

 

 

Unit-05/Lecture-03

Market-based management of Clouds (….continued)

 

Industrial implementations

Even though market-oriented models have been mostly developed in the academic domain, industrial implementations of some aspects of MOCC are becoming available and gaining popularity. In particular, some interesting initiatives show how different aspects of MOCC, such as flexible pricing models, virtual market place, and market directories, have been made available to the wider public.

 

·   Flexible pricing models: amazon spot instances

-          Amazon Web Services (AWS), one of the biggest players in the IaaS market, recently introduced the concept of spot instances, which allows EC2 customers to bid on unused Amazon EC2 capacity and run those instances for as long as their bid exceeds the current spot price.

-          The spot price varies periodically according to the supply of and demands for EC2 instances and is kept constant within a single hour block. Spot instances can be terminated at any time, and they are usually priced at a lower price with respect to the traditional (on-demand and reserved) instances, since they rely on exceeding capacity available in the EC2 infrastructure.

-           Therefore, it is the responsibility of the user to periodically persist the state of applications executing within spot instances. Spot instances represent an interesting opportunity for both Amazon and EC2 users to benefit from the current condition of the market: The provider can make revenue from a capacity that would have been wasted if priced at the normal level, and the consumer has the opportunity to pay less by taking major risks. Despite their volatile nature, spot instances have been demonstrated to be reasonably reliable and usable for performing tasks that have a lower priority and are not critical. In other words, they are suitable for applications that can tolerate QoS limitations. Moreover, they are profitably used to extend the capacity of an existing infrastructure at lower costs.

-           

·   Virtual market place: Spot Cloud

Spot Cloud is an online portal that implements a virtual marketplace, where sellers and buyers can register and trade cloud computing services. The platform is a market place operating in the IaaS sector. Buyers are looking for compute capacity that can meet the requirements of their applications, while sellers can make available their infrastructure to serve buyers ’needs and earn revenue. Spot Cloud provides a comprehensive set of features that are expected for a virtual market place. Some of them include.

 

 • Detailed logging of all the buyers ’transactions.

 • Full metering, billing for any capacity.

 • Full control over pricing and availability of capacity in the market.

 • Management of quotas and utilization levels for providers.

 • Federation management (many providers, many customers, but one platform).

 • Hybrid cloud support (internal and external resource management).

 • Full market administration and reporting.

 • Applications and pre-build appliances directories.

 

Besides being an online portal, the virtual market realized by Spot Cloud can also be replicated in the private premises. Transactions are carried out with real money and based on credit that buyers and sellers must top up once they create an account.

Spot Cloud is the most representative implementation of a platform hat enables MOCC, even though with some limitations. Spot Cloud’s working principle is a common and unique platform that sellers need to share in order to join the portal and make available computing capacity. Spot Cloud currently supports Enomaly ECP5 and Open Stack.

 

 

 

-----------------------------REFERENCE {book: M.K. Mastering of Cloud Computing, May 2013, Page No.  387-388}

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

Unit-05/Lecture-04

Federated clouds/Inter Cloud

 

Federated clouds/Inter Cloud

Cloud federation and the Inter Cloud. These are enablers for MOCC since they provide means for interoperation among different cloud providers. Cloud computing strongly implies the presence of financial agreements between parties, since services are available on demand on a pay-per-use basis. None the less, the concepts characterizing cloud federation and the Inter Cloud are applicable, with some limitations, to building aggregations of clouds that belong to different administrative domains.

 

Technology

Description

Aneka

Middleware for cloud applications development and deployment.

Broker

Middleware for scheduling distributed applications across heterogeneous systems based on the bag-of-tasks model.

Workflow management system

Middleware for the execution, composition, management, and monitoring of workflows across heterogeneous systems.

Market Maker/Meta- Broker

A matchmaker that matches the user’s requirements with service providers’ capabilities within the context of a marketplace.

InterCloud

A framework for the federation of independent computing clouds .

MetaCDN

Middleware that leverages storage clouds for intelligently delivering users’ content based on their QoS and budget preferences.

Energy-efficient computing

Ongoing research on developing techniques and technologies for addressing scalability and energy efficiency.

 

Characterization and definition

·       The terms cloud federation and InterCloud, often used interchangeably, conveys the general meaning of an aggregation of cloud computing providers that have separate administrative domains. It is important to clarify what these two terms mean and how they apply to cloud computing.

·       The term federation implies the creation of an organization that supersedes the decisional and administrative power of the single entities and that acts as a whole. Within a cloud computing context, the word federation does not have such a strong connotation but implies that there are agreements between the various cloud providers, allowing them to leverage each other’s services in a privileged manner.

·        A definition of the term cloud federation was given by Reuven Cohen, founder and CTO of Enomaly Cloud federation manages consistency and access controls when two or more independent geo-graphically distinct Clouds share either authentication, files, computing resources, command and control or access to storage resources. This definition is broad enough to include all the different expressions of cloud services aggregations that are governed by agreements between cloud providers, rather than composed by the user.

·       InterCloud is a term that is often used interchangeably to express the concept of Cloud federation. It was introduced by Cisco for expressing a composition of clouds that are interconnected by means of open standards to provide a universal environment that leverages cloud computing services.

·       By mimicking the Internet term, often referred as the “network of networks,” Inter Cloud represents a “Cloud of Clouds” and therefore expresses the same concept of federating together clouds that belong to different administrative organizations.

·       The primary difference between the Inter Cloud and federation is that the Inter Cloud is based on future standards and open interfaces, while federation uses a vendor version of the control plane. With the Inter Cloud vision, all Clouds will have a common understanding of how applications should be deployed.

·       Eventually workloads submitted to a Cloud will include enough of a definition (resources, security, service level, geo location, etc.) that the Cloud is able to process the request and deploy the application. This will create the true utility model, where all the requirements are met by the definition and the application can execute “as is “in any Cloud with the resources to support it.

·        Therefore, the term Inter Cloud refers mostly to a global vision in which interoperability among different cloud providers is governed by standards, thus creating an open platform where applications can shift workloads and freely compose services from different sources.

·       On the other hand, the concept of a cloud federation is more general and includes ad hoc aggregations between cloud providers on the basis of private agreements and proprietary interfaces.

 

-----------------------------REFERENCE {book: M.K. Mastering of Cloud Computing, May 2013, Page No.  390-392}

                                      

 

S.NO

RGPV QUESTIONS

Year

Marks

1

What kind of standards and protocols can be used to achieve interoperability in cloud federation?

Dec 2014

7

2

What are the three different levels expressing the concept of inter cloud/ cloud federation?

June 2015

7

Unit-05/Lecture-05

Cloud Federation stack

 

·         Cloud Federation stack

Creating a cloud federation involves research and development at different levels: conceptual, logical and operational, and infrastructural. Figure 5.4 provides a comprehensive view of the challenges faced in designing and implementing an organizational structure that coordinates together cloud services that belong to different administrative domains and makes them operate with in a context of a single unified service middleware. Each cloud federation level presents different challenges and operates at a different layer of the IT stack. It then requires the use of different approaches and technologies. Taken together, the solutions to the challenges faced at each of these levels constitute a reference model for a cloud federation.

 

Ř Conceptual level:

-          The conceptual level addresses the challenges in presenting a cloud federation as a favorable solution with respect to the use of services leased by single cloud providers.

-          In this level it is important to clearly identify the advantages for either service providers or service consumers in joining a federation and to delineate the new opportunities that a federated environment creates with respect to the single provider solution. Elements of concern at this level are:

 

-          • Motivations for cloud providers to join a federation

Motivations for service consumers to leverage a federation.

• Advantages for providers in leasing their services to other providers.

 • Obligations of providers once they have joined the federation

• Trust agreements between providers

 

 

Figure 5.4 Cloud federation reference stack

 

• Transparency versus consumers Among these aspects, the most relevant is the motivations of both service providers and consumers in joining a federation. From the perspective of cloud service providers, being part of federation is favorable if it helps increase their revenue and if it provides new opportunities to increase their business.

 

Moreover, the option of joining a federation can also be considered convenient if it helps sustain the QoS ensured to customers in periods of peak load, which put extreme demand on the infrastructure of the single provider. More precisely, it is possible to identify functional and nonfunctional requirements that cloud service providers have behind these motivations.

 

 The functional requirements include:

-          Supplying low-latency access to customers, regardless of their location. It is very unlikely that single cloud providers have a capillary distribution of their datacenters. Therefore, services that require low latency might provide poor performance because of unfortunate geo-location. Within this scenario the federation might help the single providers deliver the same service and meet the expected QoS.

-          Handling bursts in demand. Even though cloud computing gives the illusion of infinite capacity and continuous availability, service providers rely on a finite. IT infrastructure that eventually will be fully utilized. A natural solution to this problem is increasing the infrastructure by adding more capacity.

For example, to keep up with the increasing demand for storage and computation, Google has increased its number of servers from 8,000 to more than 450,000 in five years and moved from four server farms to more than 60 datacenters, Facebook has recently doubled its datacenter capacity. Such huge provisions are affordable for large IT companies that can make appropriate forecasts about increasing demand. Irregular demand can be better addressed by renting capacity from other providers, since not every cloud provider is in the position of being an IT giant. Cloud federation facilitates such activity by providing a context within which the lease of resources or services is encouraged.

-              Scaling existing applications and services beyond the capabilities of the owned infrastructure. The need for additional capacity can also originate from the grow thin scale of existing applications that are temporarily hosted and do not constitute a vital part of the service provider core business. Again, the opportunities for leasing additional services from a federated provider can constitute a potential advantage for a cloud federation.

-               Make revenue from unused capacity. To provide the illusion of continuous availability and infinite capacity, cloud service providers generally own large computing systems, which generate costs in terms of maintenance and power consumption despite their real use. Energy- efficient computing solutions can help reduce costs and the impact of IT on the environment. A different opportunity is given by the cloud federation, where by providers can lease their services to other providers for a limited period of time and thus make revenue, even without direct customers.

 

 

-----------------------------REFERENCE {book: M.K. Mastering of Cloud Computing, May 2013, Page No.  392-394}

 

 

 

 

 

S.NO

RGPV QUESTIONS

Year

Marks

1

Describe the architecture of cloud federation stack.

 Dec 2013

7

 

 

 

 

 

 

 

                                                    

 

 

Unit-05/Lecture-06

                                 Cloud Federation stack (….continued)

 

Logical and operational level

-         The logical and operational level of a federated cloud identifies and addresses the challenges in devising a framework that enables the aggregation of providers that belong to different administrative domains within a context of a single overlay infrastructure, which is the cloud federation.

-         At this level, policies and rules for interoperation are defined. Moreover, this is the layer at which decisions are made as to how and when to lease a service to—or to leverage a service from— another provider. The logical component defines a context in which agreements among providers are settled and services are negotiated, whereas the operational component characterizes and shapes the dynamic behavior of the federation as a result of the single providers’ choices.

-          This is the level where MOCC is implemented and realized.

-          It is important at this level to address the following challenges:

• How should a federation be represented?

• How should we model and represent a cloud service, a cloud provider, or an agreement?

• How should we define the rules and policies that allow providers to join a federation?

 • What are the mechanisms in place for settling agreements among providers?

 • What are provider’s responsibilities with respect to each other?

 • When should providers and consumers take advantage of the federation?

 • Which kinds of services are more likely to be leased or bought?

• How should we price resources that are leased, and which fraction of resources should we lease?

-     The logical and operational level provides opportunities for both academia and industry. Whereas the need for a federation—or more generally, some sort of interoperation—has now been assessed, there is no common and clear guideline for defining a model for cloud federation and addressing these challenges. Indeed, several initiatives are developing.

-     Particular attention on this level has been put on the necessity for SLAs and their definition. The need for SLAs is an accepted fact in both academy and industry, since SLAs define more clearly what is leased or bought between different providers. Moreover, SLAs allow us to assess whether the services traded are delivered according to the expected quality profile. It is then possible to specify policies that regulate the transactions among providers and establish penalties in case of degraded service delivery. This is particularly important because it increases the level of trust that each party puts in cloud federation.

-     SLAs define the provider’s performance delivery ability, the consumer’s performance profile, and the means to monitor and measure the delivered performance. An implementation of an SLA should specify:

• Purpose. Objectives to achieve by using a SLA.

• Restrictions. Necessary steps or actions that need to be taken to ensure that the requested level of service is delivered.

• Validity period. Period of time during which the SLA is valid.

• Scope. Services that will be delivered to the consumer and services that are outside the SLA.

• Parties. Any involved organizations or individual and their roles (e.g. provider, consumer).

• Service-level objectives (SLOs). Levels of services on which both parties agree. These are expressed by means of service-level indicators such as availability, performance, and reliability.

• Penalties. The penalties that will occur if the delivered service does not achieve the defined SLOs.

• Optional services. Services that are not mandatory but might be required.

 • Administration. Processes that are used to guarantee that SLOs are achieved and the related organization responsibilities for controlling these processes.

Ř Infrastructural level

- The infrastructural level addresses the technical challenges involved in enabling heterogeneous cloud computing systems to interoperate seamlessly.

- It deals with the technology barriers that keep separate cloud computing systems belonging to different administrative domains. By having standardized protocols and interfaces, these barriers can be overcome.

- In other words, this level for the federation is what the TCP/IP stack is for the Internet: a model and a reference implementation of the technologies enabling the interoperation of systems. The infrastructural level lays its foundations in the IaaS and PaaS layers of the Cloud Computing Reference Model.

- Services for interoperation and interface may also find implementation at the SaaS level, especially for the realization of negotiations and of federated clouds.

-At this level it is important to address the following issues:

 • What kind of standards should be used?

 • How should design interfaces and protocols be designed for interoperation?

 • Which are the technologies to use for interoperation?

 • How can we realize a software system, design platform components, and services enabling interoperability?

-                     Interoperation and composition among different cloud computing vendors is possible only by means of open standards and interfaces.

-                     Moreover, interfaces and protocols change considerably at each layer of the Cloud Computing Reference Model. As the more mature layer, the IaaS layer has evolved more in this sense. Almost every IaaS provider exposes Web interfaces for packaging virtual machine templates, launching, monitoring, and terminating virtual instances.

-                      Even though not standardized, these interfaces leverage the Web services and are quite similar to each other. The use of a common technology simplifies the interoperation among vendors, since a minimum amount of code is required to enable such interoperation. These APIs allow for defining an abstraction layer that uniformly accesses the services of several IaaS vendors.

-                     There are already tools—both open-source and commercial—and specifications that provide interoperability by implementing such a layer.

 

Composition is also important in considering interoperability across cloud computing platforms operating at different layers. Even in this case it is important to note that, currently, cloud computing providers operating at one layer often implement on their own infrastructure any lower layer that is required to provide the service to the end user, and they are not willing to open their stack of technologies to support interoperation. The vision proposed by a federated environment of cloud service vendors still poses a lot of challenges at each level, especially the logical and infrastructural level, where appropriate system organizations need to be designed and effective technologies need to be deployed. Considerable research effort has been carried out on the logical and operational level, and initial implementations and drafts of interoperable technologies are now developed especially for the IaaS market segment.

 

-----------------------------REFERENCE {book: M.K. Mastering of Cloud Computing, May 2013, Page No.  396-399}

 

 

Unit-05/Lecture-07

Third-party cloud services

                                                

Third-party cloud services

One of the key elements of cloud computing is the possibility of composing services that belong to different vendors or integrating them in to existing software systems. The service oriented model, which is the basis of cloud computing, facilitates such an approach and provides the opportunity for developing a new class of services that can be called third-party cloud services.

 

Examples of third-party services:

Ř MetaCDN

·         MetaCDN provides users with a Content Delivery Network (CDN) service by leveraging and harnessing together heterogeneous storage clouds.

·         It implements a software overlay that coordinates the service offerings of different cloud storage vendors and uses them as distributed elastic storage on which the user content is stored.

·          MetaCDN provides users with the high-level services of a CDN for content distribution and interacts with the low-level interfaces of storage clouds to optimally place the user content in accordance with the expected geography of its demand. By leveraging the cloud as a storage back-end it makes a complex—and generally expensive—content delivery service available to small enterprises.

·         The architecture of MetaCDN is shown in Figure 5.5

 

·         The MetaCDN interface exposes its services through users and applications through the Web; users interact with a portal, while applications take advantage of the programmatic access provided by means of Web services.

 

 

 

                    

                                     Figure 5.5 MetaCDN architecture                      

 

·         The main operations of MetaCDN are the creation of deployments over storage clouds and their management. The portal constitutes a more intuitive interface for users with basic requirements, while the Web service provides access to the full capabilities of MetaCDN and allows for more complex and sophisticated deployment.

 

·         In particular, four different deployment options can be selected:

Ř Coverage and performance-optimized deployment. In this case MetaCDN will deploy as many replicas as possible to all available locations.

Ř  Direct deployment. In this case MetaCDN allows the selection of the deployment regions for the content and will match the selected regions with the supported providers serving those areas.

Ř  Cost-optimized deployment. In this case MetaCDN deploys as many replicas in the locations identified by the deployment request. The available storage transfer allowance and budget will be used to deploy the replicas and keep them active for as long as possible.

Ř QoS optimized deployment. In this case MetaCDN selects the providers that can better match the QoS requirements attached to the deployment, such as average response time and throughput from a particular location.

 

·         A collection of components coordinate their activities in order to offer the services we described. These constitute the additional value that MetaCDN brings on top of the direct use of storage clouds by the users. Of particular importance are three components.

·         The MetaCDN Manager, the MetaCDN QoS Monitor, and the Load Redirector. The Manager is responsible for ensuring that all the content deployments are meeting the expected QoS. It is supported in this activity by the Monitor, which constantly probes storage providers and monitors data transfers to assess the performance of each provider.

 

Ř  SpotCloud

·         SpotCloud has already been introduced as an example of a virtual marketplace. By acting as an intermediary for trading compute and storage between consumers and service providers, it provides the two parties with added value.

·          For service consumers, it acts as a market directory where they can browse and compare different IaaS service offerings and select the most appropriate solution for them. For service providers it constitutes an opportunity for advertising their offerings.

·          In addition, it allows users with available computing capacity to easily turn themselves into service providers by deploying the runtime environment required by SpotCloud on their infrastructure.

 

Figure 5.6  SpotCloud market architecture

 

·         SpotCloud is not only an enabler for IaaS providers and resellers, but its intermediary role also includes a complete bookkeeping of the transactions associated with the use of resources. Users deposit credit on their SpotCloud account and capacity sellers are paid following the usual payperuse model.

·         SpotCloud retains a percentage of the amount billed to the user. Moreover, by leveraging a uniform runtime environment and virtual machine management layer, it provides users with a vendor lock-in-free solution, which might be strategic for specific applications.

 

 

·          The two previously presented examples give an idea of how different in nature third-party services can be: MetaCDN provides end users with a different service from the simple cloud storage offerings; SpotCloud does not change the type of service that is finally offered to end users, but it enriches it with additional features that result in more effective use of it.

 

 

 

 

 

-----------------------------REFERENCE {book: M.K. Mastering of Cloud Computing, May 2013, Page No.  422-426}

 

 

 

 

 

 

S.NO

RGPV QUESTIONS

Year

Marks

1

What is a third party cloud service? Give some examples of a third party cloud service.

 Dec 2014, June 2015

7

 

 

 

 

 

 

                                                    

 

 

 

 

Unit-05/Lecture-08

Case study: Google app engine and Azure

 

Google App Engine                               

 

·                              Google AppEngine is a PaaS implementation that provides services for developing and hosting scalable Web applications. AppEngine is essentially a distributed and scalable runtime environment that leverages Google’s distributed infrastructure to scale out applications facing a large number of requests by allocating more computing resources to them and balancing the load among them. The runtime is completed by a collection of services that allow developers to design and implement applications that naturally scale on AppEngine. Developers can develop applications in Java, Python, and Go, a new programming language developed by Google to simplify the development of Web applications. Application usage of Google resources and services is metered by AppEngine, which bills users when their applications finish their free quotas.

 

Architecture and core concepts

AppEngine is a platform for developing scalable applications accessible through the Web (see Figure). The platform is logically divided into four major components:

 

1.Infrastructure,

2.The runtime environment,

3.The underlying storage, and

4.The set of scalable services that can be used to develop applications

 

1.Infrastructure:

 

 AppEngine hosts Web applications, and its primary function is to serve users requests efficiently. To do so, AppEngine’s infrastructure takes advantage of many servers available within Google datacenters. For each HTTP request, AppEngine locates the servers hosting the application that processes the request, evaluates their load and, if necessary, allocates additional resources (i.e., servers) or redirects the request to an existing server. The particular design of applications, which does not expect any state information to be implicitly maintained between requests to the same application, simplifies the work of the infrastructure, which can redirect each of the requests to any of the servers hosting the target application or even allocate a new one. The infrastructure is also responsible for monitoring application performance and collecting statistics on which the billing is calculated.

 

                        

 

Figure 5.7: Google AppEngine platform architecture.

 

2.Runtime environment:

 The runtime environment represents the execution context of applications hosted on AppEngine. With reference to the AppEngine infrastructure code, which is always active and running, the runtime comes into existence when the request handler starts executing and terminates once the handler has completed.

 

·                              Sandboxing

One of the major responsibilities of the runtime environment is to provide the application environment with an isolated and protected context in which it can execute without causing a threat to the server and without being influenced by other applications. In other words, it provides applications with a sandbox.

 

Currently, AppEngine supports applications that are developed only with managed or interpreted languages, which by design require a runtime for translating their code into executable instructions. Therefore, sandboxing is achieved by means of modified runtimes for applications that disable some of the common features normally available with their default implementations. If an application tries to perform any operation that is considered potentially harmful, an exception is thrown and the execution is interrupted. Some of the operations that are not allowed in the sandbox include writing to the server’s file system; accessing computer through network besides using Mail, UrlFetch, and XMPP; executing code outside the scope of a request, a queued task, and a cron job; and processing a request for more than 30 seconds.

 

·                              Supported runtimes:

 Currently, it is possible to develop AppEngine applications using three different languages and related technologies: Java, Python, and Go.

 

AppEngine currently supports Java 6, and developers can use the common tools for Web application development in Java, such as the Java Server Pages (JSP), and the applications interact with the environment by using the Java Servlet standard. Furthermore, access to AppEngine services is provided by means of Java libraries that expose specific interfaces of provider-specific implementations of a given abstraction layer. Developers can create applications with the AppEngine Java SDK, which allows developing applications with either Java 5 or Java 6 and by using any Java library that does not exceed the restrictions imposed by the sandbox.

 

Support for Python is provided by an optimized Python 2.5.2 interpreter. As with Java, the runtime environment supports the Python standard library, but some of the modules that implement potentially harmful operations have been removed, and attempts to import such modules or to call specific methods generate exceptions. To support application development, AppEngine offers a rich set of libraries connecting applications to AppEngine services. In addition, developers can use a specific Python Web application framework, called webapp, simplifying the development of Web applications.

 

The Go runtime environment allows applications developed with the Go programming language to be hosted and executed in AppEngine. Currently the release of Go that is supported by AppEngine is r58.1. The SDK includes the compiler and the standard libraries for developing applications in Go and interfacing it with AppEngine services. As with the Python environment, some of the functionalities have been removed or generate a runtime exception. In addition, developers can include third-party libraries in their applications as long as they are implemented in pure Go.

 

3.Storage

AppEngine provides various types of storage, which operate differently depending on the volatility of the data. There are three different levels of storage: in memory-cache, storage for semistructured data, and long-term storage for static data. In this section, we describe DataStore and the use of static file servers. We cover MemCache in the application services section.

 

·                              Static file servers:

              Web applications are composed of dynamic and static data. Dynamic data are a result             of the logic of the application and the interaction with the user. Static data often are mostly constituted of the components that define the graphical layout of the application (CSS files, plain HTML files, JavaScript files, images, icons, and sound files) or data files. These files can be hosted on static file servers, since they are not frequently modified. Such servers are optimized for serving static content, and users can specify how dynamic content should be served when uploading their applications to AppEngine.

 

·                              DataStore

DataStore is a service that allows developers to store semistructured data. The service is designed to scale and optimized to quickly access data. DataStore can be considered as a large object database in which to store objects that can be retrieved by a specified key. Both the type of the key and the structure of the object can vary.

 

With respect to the traditional Web applications backed by a relational database, DataStore imposes less constraint on the regularity of the data but, at the same time, does not implement some of the features of the relational model (such as reference constraints and join operations). These design decisions originated from a careful analysis of data usage patterns for Web applications and were taken in order to obtain a more scalable and efficient data store.

 

DataStore provides high-level abstractions that simplify interaction with Bigtable. Developers define their data in terms of entity and properties, and these are persisted and maintained by the service into tables in Bigtable. An entity constitutes the level of granularity for the storage, and it identifies a collection of properties that define the data it stores. Properties are defined according to one of the several primitive types supported by the service. Each entity is associated with a key, which is either provided by the user or created automatically by AppEngine.

 

DataStore also provides facilities for creating indexes on data and to update data within the context of a transaction. Indexes are used to support and speed up queries. A query can return zero or more objects of the same kind or simply the corresponding keys. It is possible to query the data store by specifying either the key or conditions on the values of the properties. Returned result sets can be sorted by key value or properties value.

 

The implementation of transaction is limited in order to keep the store scalable and fast. AppEngine ensures that the update of a single entity is performed atomically. Multiple operations on the same entity can be performed within the context of a transaction. It is also possible to update multiple entities atomically. This is only possible if these entities belong to the same entity group. The entity group to which an entity belongs is specified at the time of entity creation and cannot be changed later. With regard to concurrency, AppEngine uses an optimistic concurrency control: If one user tries to update an entity that is already being updated, the control returns and the operation fails. Retrieving an entity never incurs into exceptions.

 

4.Application services:

 

Applications hosted on AppEngine take the most from the services made available through the runtime environment. These services simplify most of the common operations that are performed in Web applications: access to data, account management, integration of external resources, messaging and communication, image manipulation, and asynchronous computation.

 

·                              UrlFetch:

       Web 2.0 has introduced the concept of composite Web applications. Different resources are put together and organized as meshes within a single Web page. Meshes are fragments of HTML generated in different ways. They can be directly obtained from a remote server or rendered from an XML document retrieved from a Web service, or they can be rendered by the browser as the result of an embedded and remote component. A common characteristic of all these examples is the fact that the resource is not local to the server and often not even in the same administrative domain. Therefore, it is fundamental for Web applications to be able to retrieve remote resources.

 

·                              MemCache

 AppEngine provides developers with access to fast and reliable storage, which is DataStore. Despite this, the main objective of the service is to serve as a scalable and long-term storage, where data are persisted to disk redundantly in order to ensure reliability and availability of data against failures. This design poses a limit on how much faster the store can be compared to other solutions, especially for objects that are frequently accessed—for example, at each Web request.

 

AppEngine provides caching services by means of MemCache. This is a distributed in-memory cache that is optimized for fast access and provides developers with a volatile store for the objects that are frequently accessed. The caching algorithm implemented by MemCache will automatically remove the objects that are rarely accessed. The use of MemCache can significantly reduce the access time to data; developers can structure their applications so that each object is first looked up into MemCache and if there is a miss, it will be retrieved from DataStore and put into the cache for future lookups.

 

·                              Mail and instant messaging:

 Communication is another important aspect of Web applications. It is common to use email for following up with users about operations performed by the application. Email can also be used to trigger activities in Web applications. To facilitate the implementation of such tasks, AppEngine provides developers with the ability to send and receive mails through Mail. The service allows sending email on behalf of the application to specific user accounts. It is also possible to include several types of attachments and to target multiple recipients. Mail operates asynchronously, and in case of failed delivery the sending address is notified through an email detailing the error.

 

·                              Account management:

Web applications often keep various data that customize their interaction with users. These data normally go under the user profile and are attached to an account. AppEngine simplifies account management by allowing developers to leverage Google account management by means of Google Accounts. The integration with the service also allows Web applications to offload the implementation of authentication capabilities to Google’s authentication system.

 

 

 

Compute services

Web applications are mostly designed to interface applications with users by means of a ubiquitous channel, that is, the Web. Most of the interaction is performed synchronously: Users navigate the Web pages and get instantaneous feedback in response to their actions. This feedback is often the result of some computation happening on the Web application, which implements the intended logic to serve the user request. Sometimes this approach is not applicable—for example, in long computations or when some operations need to be triggered at a given point in time. A good design for these scenarios provides the user with immediate feedback and a notification once the required operation is completed. AppEngine offers additional services such as Task Queues and Cron Jobs that simplify the execution of computations that are off-bandwidth or those that cannot be performed within the timeframe of the Web request.

 

Some other information:

 

·                              Google AppEngine is a scalable runtime environment mostly devoted to executing Web applications. These take advantage of the large computing infrastructure of Google to dynamically scale as the demand varies over time.

 

·                               AppEngine provides both a secure execution environment and a collection of services that simplify the development of scalable and high-performance Web applications.

 

·                              These services include in-memory caching, scalable data store, job queues, messaging, and cron tasks. Developers can build and test applications on their own machines using the AppEngine software development kit (SDK), which replicates the production runtime environment and helps test and profile applications.

 

·                               Once development is complete, developers can easily migrate their application to AppEngine, set quotas to contain the costs generated, and make the application available to the world. The languages currently supported are Python, Java, and Go.

 

·                              Google App Engine, a cloud computing platform for hosting web application in existing Google infrastructure, it’s easy to scale, manage and free to use up to a predefined consumed resources, and it supports Java. For additional charged, please refer to this GAE billing .

 

·                              Google App Engine is a Platform as a Service (PaaS) offering that lets you build and run applications on Google’s infrastructure. App Engine applications are easy to build, easy to maintain, and easy to scale as your traffic and data storage needs change. With App Engine, there are no servers for you to maintain. You simply upload your application and it’s ready to go.

 

 

·                              Google App Engine supports apps written in a variety of programming languages.

Ř      Java: Using App Engine’s Java runtime environment, you can build your application using standard Java technologies.

Ř      Python: App Engine features a fast Python interpreter and standard Python libraries.

Ř      PHP: App Engine uses Google's Cloud Platform services under the hood when you call standard PHP functions.

Ř      Go: App Engine features a Go runtime environment that runs natively compiled Go code.

Google App Engine makes it easy to build and deploy an application that runs reliably even under heavy load and with large amounts of data. It includes the following features:

  • Persistent storage with queries, sorting, and transactions.
  • Automatic scaling and load balancing.
  • Asynchronous task queues for performing work outside the scope of a request.
  • Scheduled tasks for triggering events at specified times or regular intervals.

 

 

 

-----------------------------REFERENCE {book: M.K. Mastering of Cloud Computing, May 2013, Page No.  332-339}

 

 

 

 

 

 

S.NO

RGPV QUESTIONS

Year

Marks

1

Describe the major cloud features of Google application engine.

 Dec 2013

7

 

 

 

 

 

 

 

 

 

Microsoft Azure:

 

Microsoft Windows Azure is a cloud operating system built on top of Microsoft datacenters’ infrastructure and provides developers with a collection of services for building applications with cloud technology. Services range from compute, storage, and networking to application connectivity access control, and business intelligence. Any application that is built on the Microsoft technology can be scaled using the Azure platform, which integrates the scalability features into the common Microsoft technologies such as Microsoft Windows Server 2008, SQL Server, and ASP.NET.

 

Figure provides an overview of services provided by Azure. These services can be managed and controlled through the Windows Azure Management Portal, which acts as an administrative console for all the services offered by the Azure platform. In this section, we present the core features of the major services available with Azure.

 

Azure core concepts

The Windows Azure platform is made up of a foundation layer and a set of developer services that can be used to build scalable applications. These services cover compute, storage, networking, and identity management, which are tied together by middleware called AppFabric. This scalable computing environment is hosted within Microsoft datacenters and accessible through the Windows Azure Management Portal. Alternatively, developers can recreate a Windows Azure environment (with limited capabilities) on their own machines for development and testing purposes. In this section, we provide an overview of the Azure middleware and its services.

 

Compute services

Compute services are the core components of Microsoft Windows Azure, and they are delivered by means of the abstraction of roles. A role is a runtime environment that is customized for a specific compute task. Roles are managed by the Azure operating system and instantiated on demand in order to address surges in application demand. Currently, there are three different roles:Web role, Worker role , and Virtual Machine (VM) role.

 

·         Web role

The Web role is designed to implement scalable Web applications. Web roles represent the units of deployment of Web applications within the Azure infrastructure. They are hosted on the IIS 7 Web Server, which is a component of the infrastructure that supports Azure. When Azure detects peak loads in the request made to a given application, it instantiates multiple Web roles for that application and distributes the load among them by means of a load balancer.

                                          

 

FIGURE 5.8: Microsoft Windows Azure Platform Architecture

 

·         Worker role

              Worker roles are designed to host general compute services on Azure. They can be used to quickly provide compute power or to host services that do not communicate with the external world through HTTP. A common practice for Worker roles is to use them to provide background processing for Web applications developed with Web roles.

           

              Developing a worker role is like a developing a service. Compared to a Web role whose computation is triggered by the interaction with an HTTP client (i.e., a browser), a Worker role runs continuously from the creation of its instance until it is shut down. The Azure SDK provides developers with convenient APIs and libraries that allow connecting the role with the service provided by the runtime and easily controlling its startup as well as being notified of changes in the hosting environment. As with Web roles, the .NET technology provides complete support for Worker roles, but any technology that runs on a Windows Server stack can be used to implement its core logic. For example, Worker roles can be used to host Tomcat and serve JSP-based applications.

 

·         Virtual machine role

The Virtual Machine role allows developers to fully control the computing stack of their compute service by defining a custom image of the Windows Server 2008 R2 operating system and all the service stack required by their applications. The Virtual Machine role is based on the Windows Hyper-V virtualization technology, which is natively integrated in the Windows server technology at the base of Azure.

 

Storage services

Compute resources are equipped with local storage in the form of a directory on the local file system that can be used to temporarily store information that is useful for the current execution cycle of a role. If the role is restarted and activated on a different physical machine, this information is lost.

Windows Azure provides different types of storage solutions that complement compute services with a more durable and redundant option compared to local storage. Compared to local storage, these services can be accessed by multiple clients at the same time and from everywhere, thus becoming a general solution for storage

 

·         Blobs

Azure allows storing large amount of data in the form of binary large objects (BLOBs) by means of the blobs service. This service is optimal to store large text or binary files. Two types of blobs

are available:

 

-          Block blobs.

Block blobs are composed of blocks and are optimized for sequential access; therefore they are appropriate for media streaming. Currently, blocks are of 4 MB, and a single block blob can reach 200 GB in dimension.

 

-          Page blobs.

Page blobs are made of pages that are identified by an offset from the beginning of the blob. A page blob can be split into multiple pages or constituted of a single page. This type of blob is optimized for random access and can be used to host data different from streaming.

Currently, the maximum dimension of a page blob can be 1 TB.

 

Azure drive:

Page blobs can be used to store an entire file system in the form of a single Virtual Hard Drive (VHD) file. This can then be mounted as a part of the NTFS file system by Azure compute resources, thus providing persistent and durable storage. A page blob mounted as part of an NTFS tree is called an Azure Drive.

 

 

-----------------------------REFERENCE {book: M.K. Mastering of Cloud Computing, May 2013, Page No.  341-346}

 

 

 

 

Unit-05/Lecture-09

Case study: Hadoop, Amazon and Aneka

 

Hadoop         

 

·                  Hadoop is an open-source framework that allows to store and process big data in a distributed environment across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage.

 

·                  Hadoop runs applications using the MapReduce algorithm, where the data is processed in parallel on different CPU nodes. In short, Hadoop framework is capabale enough to develop applications capable of running on clusters of computers and they could perform complete statistical analysis for huge amounts of data.

 

 

Figure 5.9: Hadoop Framework

 

·                  Hadoop is an Apache open source framework written in java that allows distributed processing of large datasets across clusters of computers using simple programming models. A Hadoop frame-worked application works in an environment that provides distributed storage and computation across clusters of computers. Hadoop is designed to scale up from single server to thousands of machines, each offering local computation and storage.

 

 

 

Hadoop Architecture

Hadoop framework includes following four modules:

  • Hadoop Common:

These are Java libraries and utilities required by other Hadoop modules. These libraries provide file system and OS level abstractions and contains the necessary Java files and scripts required to start Hadoop.

  • Hadoop YARN: This is a framework for job scheduling and cluster resource management.
  •  Hadoop Distributed File System (HDFS™): A distributed file system that provides high-throughput access to application data.
  • Hadoop MapReduce: This is YARN-based system for parallel processing of large data sets.

We can use following diagram to depict these four components available in Hadoop framework.

 

                                    

 

Figure 5.10: Hadoop Architecture

 

MapReduce:

Hadoop MapReduce is a software framework for easily writing applications which process big amounts of data in-parallel on large clusters (thousands of nodes) of commodity hardware in a reliable, fault-tolerant manner.

The term MapReduce actually refers to the following two different tasks that Hadoop programs perform:

  • The Map Task: This is the first task, which takes input data and converts it into a set of data, where individual elements are broken down into tuples (key/value pairs).

 

  • The Reduce Task: This task takes the output from a map task as input and combines those data tuples into a smaller set of tuples. The reduce task is always performed after the map task.

 

Typically both the input and the output are stored in a file-system. The framework takes care of scheduling tasks, monitoring them and re-executes the failed tasks.

 

The MapReduce framework consists of a single master JobTracker and one slave TaskTracker per cluster-node. The master is responsible for resource management, tracking resource consumption/availability and scheduling the jobs component tasks on the slaves, monitoring them and re-executing the failed tasks. The slaves TaskTracker execute the tasks as directed by the master and provide task-status information to the master periodically.

 

The JobTracker is a single point of failure for the Hadoop MapReduce service which means if JobTracker goes down, all running jobs are halted.

 

Hadoop Distributed File System:

Hadoop can work directly with any mountable distributed file system such as Local FS, HFTP FS, S3 FS, and others, but the most common file system used by Hadoop is the Hadoop Distributed File System (HDFS).

 

The Hadoop Distributed File System (HDFS) is based on the Google File System (GFS) and provides a distributed file system that is designed to run on large clusters (thousands of computers) of small computer machines in a reliable, fault-tolerant manner.

 

HDFS uses a master/slave architecture where master consists of a single NameNode that manages the file system metadata and one or more slave DataNodes that store the actual data.

A file in an HDFS namespace is split into several blocks and those blocks are stored in a set of DataNodes. The NameNode determines the mapping of blocks to the DataNodes. The DataNodes takes care of read and write operation with the file system. They also take care of block creation, deletion and replication based on instruction given by NameNode.

 

HDFS provides a shell like any other file system and a list of commands are available to interact with the file system. These shell commands will be covered in a separate chapter along with appropriate examples.

 

                 

 

Advantages of Hadoop

  • Hadoop framework allows the user to quickly write and test distributed systems. It is efficient, and it automatic distributes the data and work across the machines and in turn, utilizes the underlying parallelism of the CPU cores.
  • Hadoop does not rely on hardware to provide fault-tolerance and high availability (FTHA), rather Hadoop library itself has been designed to detect and handle failures at the application layer.
  • Servers can be added or removed from the cluster dynamically and Hadoop continues to operate without interruption.
  • Another big advantage of Hadoop is that apart from being open source, it is compatible on all the platforms since it is Java based.

 

 

 

-----------------------REFERENCE {Internetlink:http://www.tutorialspoint.com/hadoop/hadoop_introduction.htm}

 

 

 

 

S.NO

RGPV QUESTIONS

Year

Marks

1

Explain the architecture of HADOOP with its diagram.

 Dec 2013

7

 

 

 

 

Amazon:

Amazon Web Services (AWS), a collection of remote computing services, also called web services, make up a cloud-computing platform offered by Amazon.com. These services operate from 11 geographical regions across the world. The most central and well-known of these services arguably include Amazon Elastic Compute Cloud and Amazon S3. Amazon markets these products as a service to provide large computing-capacity more quickly and more cheaply than a client company building an actual physical server farm.

 

AWS offers comprehensive cloud IaaS services ranging from virtual compute, storage, and networking to complete computing stacks. AWS is mostly known for its compute and storage-on- demand services, namely Elastic Compute Cloud (EC2) and Simple Storage Service (S3).

 

EC2 provides users with customizable virtual hardware that can be used as the base infrastructure for deploying computing systems on the cloud.

 

It is possible to choose from a large variety of virtual hardware configurations, including GPU and cluster instances. EC2 instances are deployed either by using the AWS console, which is a comprehensive Web portal for accessing AWS services, or by using the Web services API available for several programming languages.

 

EC2 also provides the capability to save a specific running instance as an image, thus allowing users to create their own templates for deploying systems. These templates are stored into S3 that delivers persistent storage on demand. S3 is organized into buckets; these are containers of objects that are stored in binary form and can be enriched with attributes. Users can store objects of any size, from simple files to entire disk images, and have them accessible from everywhere. 

 

Besides EC2 and S3, a wide range of services can be leveraged to build virtual computing sys- tems. Including networking support, caching systems, DNS, database (relational and not) support, and others.

 

 

Why Amazon Web Services

Amazon.com initiated the evaluation of Amazon S3 for economic and performance improvements related to data backup. As part of that evaluation, they considered security, availability, and performance aspects of Amazon S3 backups. Amazon.com also executed a cost-benefit analysis to ensure that a migration to Amazon S3 would be financially worthwhile. That cost benefit analysis included the following elements:

  • Performance advantage and cost competitiveness. It was important that the overall costs of the backups did not increase. At the same time, Amazon.com required faster backup and recovery performance. The time and effort required for backup and for recovery operations proved to be a significant improvement over tape, with restoring from Amazon S3 running from two to twelve times faster than a similar restore from tape. Amazon.com required any new backup medium to provide improved performance while maintaining or reducing overall costs. Backing up to on-premises disk based storage would have improved performance, but missed on cost competitiveness. Amazon S3 Cloud based storage met both criteria.
  • Greater durability and availability. Amazon S3 is designed to provide 99.999999999% durability and 99.99% availability of objects over a given year. Amazon.com compared these figures with those observed from their tape infrastructure, and determined that Amazon S3 offered significant improvement.
  • Less operational friction. Amazon.com DBAs had to evaluate whether Amazon S3 backups would be viable for their database backups. They determined that using Amazon S3 for backups was easy to implement because it worked seamlessly with Oracle RMAN.
  • Strong data security. Amazon.com found that AWS met all of their requirements for physical security, security accreditations, and security processes, protecting data in flight, data at rest, and utilizing suitable encryption standards.

 

 

The Benefits

With the migration to Amazon S3 well along the way to completion, Amazon.com has realized several benefits, including:

 

  • Elimination of complex and time-consuming tape capacity planning. Amazon.com is growing larger and more dynamic each year, both organically and as a result of acquisitions. AWS has enabled Amazon.com to keep pace with this rapid expansion, and to do so seamlessly. Historically, Amazon.com business groups have had to write annual backup plans, quantifying the amount of tape storage that they plan to use for the year and the frequency with which they will use the tape resources. These plans are then used to charge each organization for their tape usage, spreading the cost among many teams. With Amazon S3, teams simply pay for what they use, and are billed for their usage as they go. There are virtually no upper limits as to how much data can be stored in Amazon S3, and so there are no worries about running out of resources. For teams adopting Amazon S3 backups, the need for formal planning has been all but eliminated.

 

 

  • Reduced capital expenditures. Amazon.com no longer needs to acquire tape robots, tape drives, tape inventory, data center space, networking gear, enterprise backup software, or predict future tape consumption. This eliminates the burden of budgeting for capital equipment well in advance as well as the capital expense.

 

  • Immediate availability of data for restoring – no need to locate or retrieve physical tapes. Whenever a DBA needs to restore data from tape, they face delays. The tape backup software needs to read the tape catalog to find the correct files to restore, locate the correct tape, mount the tape, and read the data from it. In almost all cases the data is spread across multiple tapes, resulting in further delays. This, combined with contention for tape drives resulting from multiple users’ tape requests, slows the process down even more. This is especially severe during critical events such as a data center outage, when many databases must be restored simultaneously and as soon as possible. None of these problems occur with Amazon S3. Data restores can begin immediately, with no waiting or tape queuing – and that means the database can be recovered much faster.

 

  • Backing up a database to Amazon S3 can be two to twelve times faster than with tape drives. As one example, in a benchmark test a DBA was able to restore 3.8 terabytes in 2.5 hours over gigabit Ethernet. This amounts to 25 gigabytes per minute, or 422MB per second. In addition, since Amazon.com uses RMAN data compression, the effective restore rate was 3.37 gigabytes per second. This 2.5 hours compares to, conservatively, 10-15 hours that would be required to restore from tape.

 

  • Easy implementation of Oracle RMAN backups to Amazon S3. The DBAs found it easy to start backing up their databases to Amazon S3. Directing Oracle RMAN backups to Amazon S3 requires only a configuration of the Oracle Secure Backup Cloud (SBC) module. The effort required to configure the Oracle SBC module amounted to an hour or less per database. After this one-time setup, the database backups were transparently redirected to Amazon S3.

 

  • Durable data storage provided by Amazon S3, which is designed for 11 nines durability. On occasion, Amazon.com has experienced hardware failures with tape infrastructure – tapes that break, tape drives that fail, and robotic components that fail. Sometimes this happens when a DBA is trying to restore a database, and dramatically increases the mean time to recover (MTTR). With the durability and availability of Amazon S3, these issues are no longer a concern.

 

  • Freeing up valuable human resources. With tape infrastructure, Amazon.com had to seek out engineers who were experienced with very large tape backup installations – a specialized, vendor-specific skill set that is difficult to find. They also needed to hire data center technicians and dedicate them to problem-solving and troubleshooting hardware issues – replacing drives, shuffling tapes around, shipping and tracking tapes, and so on. Amazon S3 allowed them to free up these specialists from day-to-day operations so that they can work on more valuable, business-critical engineering tasks.

 

  • Elimination of physical tape transport to off-site location. Any company that has been storing Oracle backup data offsite should take a hard look at the costs involved in transporting, securing and storing their tapes offsite – these costs can be reduced or possibly eliminated by storing the data in Amazon S3.

 

As the world’s largest online retailer, Amazon.com continuously innovates in order to provide improved customer experience and offer products at the lowest possible prices. One such innovation has been to replace tape with Amazon S3 storage for database backups. This innovation is one that can be easily replicated by other organizations that back up their Oracle databases to tape.

 

 

 

 

--------------------REFERENCE {Internet link: https://aws.amazon.com/solutions/case-studies/amazon/}

 

 

 

 

 

 

 

 

Aneka: A Cloud Application Platform:

 

·                  Aneka is a market-oriented cloud development and management platform with rapid application development and workload distribution capabilities. Aneka is an integrated middleware package which allows you to seamlessly build and manage an interconnected network in addition to accelerating development, deployment and management of distributed applications using Microsoft .NET frameworks on these networks. It is market-oriented since it allows you to build, schedule, provision and monitor results using pricing, accounting, QoS/SLA services in private and/or public (leased) network environments.

 

·                  Aneka is an Application Platform-as-a-Service (Aneka PaaS) for Cloud Computing. It acts as a framework for building customized applications and deploying them on either public or private Clouds. One of the key features of Aneka is its support for provisioning resources on different public Cloud providers such as Amazon EC2, Windows Azure and GoGrid.

                                                                    

·                  Aneka is a .NET-based application development Platform-as–a-Service (PaaS), which offers a runtime environment and a set of APIs that enable developers to build customized applications by using multiple programming models such as Task Programming, Thread Programming and MapReduce Programming, which can leverage the compute resources on either public or private Clouds.

 

·                  Moreover, Aneka provides a number of services that allow users to control, auto-scale, reserve, monitor and bill users for the resources used by their applications. One of key characteristics of Aneka PaaS is to support provisioning of resources on public Clouds such as Windows Azure, Amazon EC2, and GoGrid, while also harnessing private Cloud resources ranging from desktops and clusters, to virtual datacentres when needed to boost the performance of applications, as shown in Figure 5.11. Aneka has successfully been used in several industry segments and application scenarios to meet their rapidly growing computing demands.

 

 

Overview of Aneka Cloud Application Development Platform

 

Figure 5.12 shows the basic architecture of Aneka. The system includes four key components, including Aneka Master, Aneka Worker, Aneka Management Console, and Aneka Client Libraries

 

 

 

 

 

 

Figure 5.11: Aneka Cloud Application Development Platform.

 

 

-                               The Aneka Master and Aneka Worker are both Aneka Containers which represents the basic deployment unit of Aneka based Clouds. Aneka Containers host different 4 kinds of services depending on their role. For instance, in addition to mandatory services, the Master runs the Scheduling, Accounting, Reporting, Reservation, Provisioning, and Storage services, while the Workers run execution services.

-                                For scalability reasons, some of these services can be hosted on separate Containers with different roles. For example, it is ideal to deploy a Storage Container for hosting the Storage service, which is responsible for managing the storage and transfer of files within the Aneka Cloud.

-                               The Master Container is responsible for managing the entire Aneka Cloud, coordinating the execution of applications by dispatching the collection of work units to the compute nodes, whilst the Worker Container is in charge of executing the work units, monitoring the execution, and collecting and forwarding the results.

 

 

                

 

 

Figure 5.12: Basic Architecture of Aneka.

 

The Management Studio and client libraries help in managing the Aneka Cloud and developing applications that utilize resources on Aneka Cloud. The Management Studio is an administrative console that is used to configure Aneka Clouds; install, start or stop Containers; setup user accounts and permissions for accessing Cloud resources; and access monitoring and billing information. The Aneka client libraries, are Application Programming Interfaces (APIs) used to develop applications which can be executed on the Aneka Cloud. Three different kinds of Cloud programming models are available for the Aneka PaaS to cover different application scenarios:: Task Programming, Thread Programming and MapReduce Programming These models represent common abstractions in distributed and parallel computing and provide developers with familiar abstractions to design and implement applications.

 

 

 

 

--------------------REFERENCE {Internet link: http://gridbus.cs.mu.oz.au/papers/Aneka-AzurePlatform.pdf}

 

 

 

S.NO

RGPV QUESTIONS

Year

Marks

1

Write a brief note on Aneka and Hadoop.

 June 2015

7