Showing posts with label netflixOSS. Show all posts
Showing posts with label netflixOSS. Show all posts

Friday, April 27, 2018

Generic and Script Remediation

Some time ago I was sharing my experiences with a Dynomite Remediation Process I wrote. Iḿ using remediation systems for a while and IMHO the add lots of value since they automate manual Cloud Operation work and save time for people. Currently, I refactored my Remediation code and now the same code can support Dynomite / Dynomite Manager but also Apache Cassandra. There are very similar concepts between Dynomite and Cassandra Remediation such as discovering AWS EC2 Ips to be remediated, AWs Resourcing(Creating SGs, Deleting LCs, Updating ASGs), Health checking(Is the node up and running? Could I remediate right now or you are in the middle of a backup or just booting up?). Once I identified that core concept I was able to create a high level and generic design for Cassandra and Dynomite. This is great for cases of Patch / Fix(apply a new AMI) or Scale Up(Increase the memory or CPU for instance). Although that the main use cases and most important ones there were some corner cases and always it will have URGENT fixes that don't and can't wait for a new AMI. Remediation allows AMI rollback and therefore is great for immutable infrastructure. I respect and agree with Immutable Infrastructure principle however there are use cases where we need to do things differently. Don't get wrong, if you don't craft your use case with care is easy to reach a scenario where you end up harassing immutable infrastructure and therefore creating lots of side effects. That's not my goal here. However, the Hell is full of good intentions some a tool like this need to be used with lots of care otherwise bad things will happen. Right now you must be wondering why do Script remediations? Why don't use something like Terraform or Ansible and that's it? Ansible is a great tool and Terraform too, however, there are several kinds of problems and solutions. My problem is quite unique(considering my customer, my current problems, and current stack). I do believe most of the things I'm sharing here can be used by a broader audience, however, IF you don't have this problems or value DevOps Automation somethings here might sound too much.

Terraform, why not? 

First of all, Terraform is awesome. However, I have a very unique kind of problem here. I use NetflixOSS Dynomite and Dynomite Manager.Dynomite is written in C and Dynomite Manager in Java. DM(Dynomite manager) does all kind of automation work like security groups, backups and restores to S3 using Java. So I don't have an Ansible or Terraform Script to change, instead, I do have Java code which is better IMHO. If you are curious about this DevOps style of work please read this. Secondly, the kind of problem we have here(Data Layer - not microservices) required very dynamic work and this kind of work could be done by Terraform however it would require write a plugin in Go or generate Terraform code. Since My Stack(based on NetflixOSS Stack) is Java and I work on a Java shop it does make sense to keep it in Java.

The Issue with Ansible

Ansible is cool. However, it falls on the same issues like Terraform. I would need to write a plugin in Python or generate ansible code which would not be ideal. So Again being in Java would be better, since DM is in Java already and I do deliver(A guy with Engineering background) DevOps with Code(For data problems) is soo much better. In regards to ansible I have a second issue - In my current project, we don't use Ansible Inventory. We just use Ansible for provisioning, so this also makes discovery hard. Even If I had inventory it would have others problems like Microservices developers often don't like do anything related to DevOps work so they might not have even access to AWS console so asking they to use Ansible might be not ideal. IMHO developers should be responsible to DevOps as well however I think core and platform teams need to provide better tools to make that work easy.

The case for Script Remediation

The Script generation is generic. It's possible to apply bash scripts in Cassandra and Dynomite clusters using a generic Jenkins job. The Script remediation will receive a bash script which will be encoded using base64 in order to be passed as java properties from Jenkins to gradle and also a timeout. Timeout is required because the generic script remediation will connect on each cluster node and apply the script. Some scripts might take more time them other so that's why we need to pass the timeout.


The Generic Script Remediation is just a simple java code that connects on each node and applies the script. There is some information that is captured to generate a final report such as host, AZ, time to execute the script, it was successful or if there were errors. That'st it.

This code just works fine for Dynomite and Cassandra. There are some use cases for this approach in my project - like:

  • Data Mugling: Sometimes developers need to clean up data(Stress test or Investigation).
  • Ops Emergency: Security Patch, Telemetry or quick rollbacks.
Most of times Patch / Fix remediation will be used and therefore force immutable infrastructure but for that 2 use cases I described Script remediation is better and also faster. Regular(Patch/Fix) remediation takes lots of time(~16 minutes for 3 node cluster). Script remediation can be done in less than a minute. 




Since you can pass a bash script that could be a backdoor to hurt immutable infrastructure or even destroy the cluster, loose data or create availability issues or even downtimes. So this need to be used with care and lots of caution. Using with wisdom this is a great tool to save people time and reduce manual work and avoid repetitive tasks such as cleaning up data before the stress test.

Whats Next

Currently, I'm wondering If I should have a 3rd kind of remediation based on security groups because time to time you need to open ports. Since DM generate all security groups(there is no script to change it) and create java code for that is too much. However I'm not sure if that would be a real abstraction and real value or would be just another interface since I don't see developers opening ports, this is a task for a cloud ops team and they don't have issues with ansible nor terraform. So they could go to aws console or run ansible from a bastion node. Thinking this way the 3rd kind of remediation based on SG might be just another UI(Jenkins) and that where I'm. Also could not find a Java API for Terraform. I also don't want to create a new DSL  for remdiation would be better leverage Terraform syntax but them this would require to colde a plugin or generate Terraform code. So for now I will wait and see if this business case become more important, right now that did not happen a lot of times and we need be lean, automation is great but there is a cost associated with and sometyhimes is better leverage new features.

cheers,
Diego Pacheco

Friday, April 13, 2018

Dynomite Remediation

Working with the cloud is great but it's not easy. The Cloud allows us spins machines pretty easily, however, update these machines are not that an easy task, especially if we talking about databases. Today I want to share some experience with remediation process I've built for Dynomite in my current project.  I'm using NetflixOSS Dynomite in production(AWS) for more than 2 years now both as cache but also as Source of Truth. I need to say that is rock solid and just works. However often we need to increase some machine capacity, i.e: Add more memory or increase the disk size. There are times where we need to apply patches on the OS - I use Amazon Linux on Production on EC2 and recently there was the Meltdown and Spectre situation. Other times there are telemetry or even provisioning and configuration bugs that need to be fixed.  No matter the reason you will always need, time to time, change something on the underlying AMI or in the Java or C application which is running. For this times there is a need to apply changes in a very smooth way, this is hard, It's not only hard in the sense that bad things will happen if you do it wrong but also because it's a long process and it's easy to do something wrong.


Why Manual Remediation is a bad idea?

First of all, manual operations are complicated at Cloud / Microservices era, since we have more software and this software is often way more distributed it means we have more machines to deal with. More machines mean more steps because sometimes the changes are not only on the machine but on the underlying infrastructure like Launch configurations or auto-scaling groups. Apply this changes to Database infrastructure requires very specific ordering. Ordering is easy to determine but proper execution is hard. As humans we are great but we always make mistakes, also like Google SRE folks said in the past, everytime you add a human you will introduce latency. Latency might be okay for 1 cluster, but for several clusters, it's a bigger problem.  Wrapping up: Cloud Operation Manual work has these potential problems:

  • Error-Prone: Mistakes or lack of attention to detail
  • Latency: It's will take more time to get the job done
  • Scalability Issues: It will be a bottleneck if a team need to do for all customers
  • Data Loss: If some task is executed out of order you might lose data.
  • Downtime: Depending on the mistake or error it could affect availability.

Dynomite Use Cases and Topology

I use dynomite as cache but also as a source of truth. When you think as a cache only remediation is great but less critical since something went wrong or takes more time it's fine since is a cache and you are not losing any data. Worst case you are adding more latency to the application, depending on your volume that might or might not be acceptable at all.

When we are talking about Source of Truth use case is a different deal since we can't afford losing Data. Thanks to Netflix awesome engineers there is a killer feature in Dynomite Manager called Cold Bootstrapping. This feature is very important for 2 reasons, first because allow dynomite / redis nodes to a recovery process when a node dies or got terminated - so the node can peer up with other nodes and get nearby data and they join the cluster without losing data. The same feature allows to perform remediation without downtime if we are careful and execute steps in very narrow and specific order.

Typical a dynomite cluster has 3 nodes as minimum topology. There is 1 node per AZ so this helps us to achieve high availability. You can have bitter clusters but will be always multiple of 3, like 3,6,9,12,15,18,21 nodes and so on and on.



Dynomite Requirements

In order to do a proper remediation, there are 3 main requirements we need to meet. First of all, we need to be fully automated them we also need to make sure we don't lose data and we don't create downtime for our consumers.

It's possible to avoid downtime because dynomite has dyno. Dyno is cluster and topology aware so it knows other nodes in other AZ so when a node is down it will failover to the other node and that's the reason why we can do the process without downtime. Dyno has some other interesting concept called LOCAL_RACK or LOCAL_AZ which allow you to have a preferable AZ to connect this allow you reduce latency and get nearby AZ.

In order to the remediation process work, dyno needs to be fully and properly configured and you need to have a proper cluster, in sense of topology. Othwerise you might not meet the requirement of no downtime.  Automation can be archived with AWS SDK which is available to most of the popular languages. The remediation solution I built was coded with Java 8. Overall JVM ecosystem is really great for troubleshooting, remote debug and profiling so this was one of the reasons behind the rationale plus my company I work for is a Java shop so :D.


Dynomite Remediation Process Flow

Right now we can talk about the remediation process. This process had to be coded very carefully in sense of timeouts, retry and make sure the steps and execute at the right moment. Forst instance first thinks the process does is to discover the cluster based on the ASG Name. Here I use the AWS API in order to get all dynomite nodes given cluster name which is unique.

Once I have the IPs that I can check is the nodes are healthy and if is a good time to start a remediation process. The first question makes sense and is easy to grasp because we want to know if the node is UP and RUNNING howe the er second question makes more sense if we think about Dynomite-Manager(DM). DM might be running some Backup yo S3 or might be performing some cold bootstrapping with redis since a node might just die or could be restoring data from S3 no matter the case this is pretty good examples of things that could be happing and you supposed to wait for all these events finish before remediation starts or proceed. That's why calling the Dynomite-Manager Health Checker is so important and thanks to Netflix great engineer it exposes all information we need to know in order to perform this process in a safe and reliable manner.

In order to apply patches or increase instance family, we need to create new Launch configurations pointing to new AMI IDS and we need to update Auto-Scaling Groups to point to these new LCs. Them we can kill node by node in order to remediate the cluster. After all, nodes are remediated we can delete the 3 old LCs(1 per AZ) and we are done.

When we kill a node there are operations that will happen and we will need to wait for this operations to be finished in order to move to the next node. First we need wait AWS ASG bootup a new node, then we need wait dynomite and DM bootup and DM will realize it needs to run some cold bootstrap in order to copy data from other Redis nodes, once that process is done DM health Checker will let us know and then we can proceed to next node.

Here is the visual picture of the whole remediation process flow.






















Remediation process is great and saves lots of times. The whole remediation process is a bit slow since we need wait AWS to kill and boot up new instances. On average a 3 node cluster takes about 15 minutes to be remediated. This time may vary depending on how much data do you have in Redis.

Other Dynomite related Posts
Cheers,
Diego Pacheco

Saturday, November 11, 2017

Running Dynomite on AWS with Docker in multi-host network Overlay

Dynomite is a kick-ass project. Basically, allow you to have strong consistency on top of NoSQL Databases. I've been using dynomite for a while in production(AWS) and I can say the core is rock solid and it just works.

Lots of developers use Windows or Mac for instance and dynomite is built in C and it's really meant for Linux(Like all good things).  So some time ago I made 2 simple projects to get started quickly with dynomite.  Basically, the project creates a simple dynomite 3 node cluster and let you run on your local machine with docker.

There are 2 projects - One to create a dynomite cluster with Redis -- The other with Facebook's RocksDB(Experimental). So you can use it on your local machine to Debug and it works just fine. So why not go 1 step further and run Dynomite in AWS using docker? There are cool benefits if you do this approach.

Running Dynomite on AWS with Docker

Dynomite works very well on AWS but also in any other cloud-vendor or Bare metal DC. Dynomite runs on Docker just fine too. Now you can choose to run on EC2, ECS, Kubernetes on EC2 or even EC2 with docker.

The Benefits

There are many advantages do run Dynomite with docker on aws.

Here are some Benefits -- The good things:
 - COST Savings: Since you can benefit from your reservation and do better resource utilization.
 - Less Latency: Running on docker allow you to easily deploy on the same box as application and reduce network roundtrips.
 - Portability: Same docker image can be used to run anywhere also from the developer machine.
 
The Cons

Like everything in life, there are pros and cons. Here are some I found:

- Networking: Docker networking can get very tricky and hard to maintain.
- Size Limitations: Default network in /24 so it's limited to 256 ips. Offcourse you can create more networks.
- More Complex: You will have docker, docker cluster(swarm), docker network(overlay) to managed so there are more moving points of failure compared with just running dynomite on EC2 for instance.

Getting Started 

Now we will install Docker, Docker Swarm, Configure a docker cluster, Create a network overlay and run dynomite in a cluster in Docker on Ec2. Phew! Long list. :-)



We will do something very silly and simple. So will deploy a 3 node cluster. This cluster won't have sharding(You can have sharding on dynomite - just dependents on seeds config - for sake of simplicity we will not do it) or cold bootstrapping or S3 backups - If you are interested in this feature you should take a look in Dynomite-Manager.

Basically, we need do the following steps in order to get this working. These are the steps:
 1.  Create EC2 instances(Let's say 2) - Later you can automate(Ansible, Boto3, Terraform, whatever)
 2.  Create Security Groups(Use the same SG for all ec2 instances) like sg_dynomite_docker.
 3. You need open ports(SG): 8101, 8102, 6379, 2377 and any others your app might need.
 4. Them we ssh to the box and install Docker
 5. Install docker Swarm - become a master - Docker will give you the command.
 6. Do ssh to the other box and they join the master swarm node.
 7. Create a Docker network with overlay - make sure it's attachable.
 8. Configure dynomite YAML files to use fixed ips on the docker network overlay.
 9. Do docker run and run docker dynomite container 2x in 1 host
10. Do docker run and run docker dynomite on another host. That's it.

You can use my dynomite-docker project as a starting point and make the changes there because there are configs and Dockerfile done you just need change the IPs and remove the volume mapping and make sure you create the docker network with overlay as that's it. Here there is a https://gist.github.com/diegopacheco/6c75a445337e1ac29fd9ae07a16e2500 sample snipper that might help you.

Cheers,
Diego Pacheco

Friday, November 10, 2017

Building Effective Microservices - 80% OFF until 30th November

Want to learn how to build microservices using Java 8, NetflixOSS Stack(Eureka, RxNetty, Feign, Hystrix, Ribbon) using Kubernetes(Minikube). 

That's the real deal you read it right. Get my videos series: Building Effective Microservices 80% until 30th November 2017.

Saturday, November 4, 2017

Dynamic Configurations with Annotations and NetflixOSS Archaius 2

NetflixOSS Archius 2 is a great Dynamic Configuration solution for microservices. Archaius is based on Apache Commons Configurations project.

Using Archaius we can load configurations from several sources such as OS env vars or any Database like Oracle or even from Zookeeper. If there is a missing configuration source you can add it pretty easy and load your configs. 

Archaius can be used in any java project no matter if is a microservice or not. 

Archaius also support dynamic configuration refresh via callbacks -- In short, this means you can reload your configs without re-deploy or downtime in your microservice -- which is really great.

Archaius has many nice features. Archaius also is very modular and easy to extend there are some nice community extensions available as well. 



Archaius 2 has some many nice and desirable features such as:
  • Dynamic, Typed Properties
  • High throughput and Thread Safe Configuration operations
  • A polling framework that allows obtaining property changes of a Configuration Source
  • A Callback mechanism that gets invoked on effective/"winning" property mutations (in the ordered hierarchy of Configurations)
  • A JMX MBean that can be accessed via JConsole to inspect and invoke operations on properties
  • Out of the box, Composite Configurations (With ordered hierarchy) for applications (and most web applications willing to use convention based property file locations)
  • Implementations of dynamic configuration sources for URLs, JDBC and Amazon DynamoDB
  • Scala dynamic property wrappers
Using Archaius 2 with Java Annotations

build.gralde


Already so we defined all archaius 2 dependencies needed. Now we can define the configurations we need - We will do that using Java annotations.

Let's create a Java interface called AppConfig.java.


First of all, we can see a class level annotation called @Configuration with this annotation we will define a configuration prefix so in this way, we don't need to repeat our selfs in other annotations.

Here we defined 3 methods using @PropertyName annotation and we also set up some default values using @DefaultValue annotation from archaius 2.

Now we need define some DefaultPropertyFactory just to make Archaius 2 injections happy - we could apply some customizations here but we really don't need it.

DefaultPropertyFactory.java

Next, we need to define GuiceArchaiusModule where we will add all other Guice bindings and archaius 2 configurations.



Here we have a couple of configurations. We need to define a Decoder and PropertyFactory in order to use Archaius. It's not likely you need to configure this -  just do the bindings like I did.

The getconfigs method is quite interesting and cool. Basically, I'm providing configuration from multiple sources with custom code. I'm getting all system config and also reading configs from a property file called app.properties if present and also doing some trick with os env variables. I'm converting all "_" to ".". Why? Because often OS Env vars cannot have "." and in Java, we use to use "." on property names so this conversation makes OS envs compatible with java.  Also, a very simple way to customize configurations.

Ok. Now we can create a simple class to read archaius configs.

SimpleService.java


Guice will create a proxy and inject the configs for us using the config source if they are not found the default values will kick in and be used.

Finally, we have our main class in other to load the Guice and Archaius 2 modules and run the application.

Main.java


src/main/resources/app.properties
app.has.persistence=false
#app.name.simple=HA

Archaius 2 is very useful and using annotations it's very sexy and works very well for Java developers.

You can also download the code in my github here.

Cheers,
Diego Pacheco

Getting started with NetflixOSS Governator

NetflixOSS Governator is a set of Google Guice extensions to create REST services using Jersey.

Using Governator we can easily configure servers like Jetty and Tomcat in order to build microservices. We also can use set of guice modules to integrated with Archaius and Eureka-Client.

Governator is not opinionated, it's similar to Spring Boot in comparison. However, Governator is configured to work with Guice and not Spring framework.

Governator it's cool because you can define pretty much everything using java code and annotations in a declarative fashion. All code is configured in Guice so we can take benefit of Ioc and Dependency injection and end up creating solution more testable by nature.


Governator Features:
  • Classpath scanning
  • Automatic binding
  • Lifecycle management
  • Configuration of field mapping
  • Field validation
  • Parallelized object warmup

Creating a Simple REST Service with Governator 1.x

First of all, we need to define our dependencies on build.gradle. So let's defined the dependencies.

build.gradle

 

Now we can do the java code -- which is pretty easy. 

JerseyMain.java

So here we have the following. We have a simple REST resource called SimpleResource which will take HTTP request on "/" address. Them we just have to configure Governator / Guice bindings and that's it.

We need to add the ShutdownHookModule in order to have a shutdown port for our server. I'm also overriding the JettyModule in order to change the server port to 9090. Them we need to add the Governator support for Jersey with GovernatorJerseySupportModule. Where we configure the PREFIX of all REST resource, in this case, will be "/*". We also are binding all REST resource with getResourceConfig. That's it we have a working service with we can run with $ ./gradle run or just run as the main class in eclipse for instance.

You can download the code source from my Github here.

Cheers,
Diego Pacheco

Friday, September 29, 2017

Dynomite Eureka Registry with Prana

Dynomite it's a great solution for clustering with NoSQL Databases like Redis. Eureka is a nice Registry & Discoverability Solution.

We can get best of both worlds using Prana. Prana is a sidecar that enables non-JVM applications to register in Eureka.

Sometimes we could easily use multiple discoverability solutions like DNS, ETCD, Eureka etc... However not all discoverability provide the same benefits some tools are better suited for some jobs them other. I like ETCD and make sense onKubernetes world but if you are doing Java Microservices eureka makes more sense.

The triad(Eureka, Prana, Dynomite) is great, This is great because then you can do discoverability on your database nodes, this is not great for several reasons like:
  • Use the same tool for Registry / Discoverability
  • Enable all sorts of dynamic programming which is great for DevOps Engineering
  • Avoid AWS Throttling issues
  • Make dyno clients more dynamic and this is a better solution them DNS like route53
I made a simple video with a simple presentation and live demo how to do this work on the server side with Dynomite 0.5.9. I hope you enjoy, have fun.


Slides from the presentation



Update: After talking with a Netflix Engineer, I got some new information about the state of Prana which can be found here. If you are using Dynomite-manager you should let DM do the Registry for you so there is no need to use Prana. For DM case you check this out. IF you are using Dynomite without DM you can still use Prana however you would consider DM in case you are running on AWS.

Cheers,
Diego Pacheco

Saturday, June 25, 2016

Running Netflix Dynomite and Dynomite-Manager at AWS Cloud

Netflix Dynomite is Kick Ass Generic Dynamo implementation for K/V Stores. Recently Netflix released the Dynomite-Manager at he NetflixOSS Meetup Season 4 Episode 2. Dynomite and Dynomite-Manager are Rock Solid both are a beautiful piece of engineering work.

Dynomite has High Throughput and Low Latency. It gives superpowers to Redis making him Strong Consistent with Quorum like semantics and multi-datacenter.  It's possible to use Dynomite as Cache or as a Data Store.


Features

Dynomite-Manager has neat features like:

  • Cold Bootstrapping
  • S3 Backups and Restores
  • Cluster Management via REST Apis
  • Token Management 
NetflixOSS Dynomite and Dynomite-Manager Presentation


Getting Started

You can get more details about Dynomite-Manager on this blog post. In this blog post, I will show how to Install Dynomite and Dynomite-Manager step-by-step in the AWS EC2 Cloud. Beware this is reference post you should not use these configs in your production Env. However, this post should help you to get Dynomite Manager up and running :D. So Let's get Started!



Cheers,
Diego Pacheco



Monday, May 16, 2016

@NetflixOSS meetup: Season 4 Episode 2

June 1st, 2016 I will be speaking about my experiences with NetflixOSS and Dynomite at the Season 4 Episode 2 Netflix meetup.

You can check out some posts I did about the NetflixOSS stack:

There are POCs and code samples available in my Github which you can check it out here. 

Cheers,
Diego Pacheco

Tuesday, November 10, 2015

Netflix Dynomite/Dyno: The Cluster for Redis


Dynomite is brilliant. Kudos for NetflixOSS team because it kicks ass. First all they mixed several interesting, battle tested and sexy architectural ideas and deliver into a single solution. What would be Dynomite? You can think as a kick Ass Cluster for Memcached and Redis. But its way more than that. Dynomite is integrated with the Netflix Stack so you can use with Eureka and the rest of the stack. You dont need use Redis or Memcached if you dont want because Dynomite is modular so you can use the NoSQL or thing behind it.

Dynomite is based on the Amazon Dynamo paper, so it implements the Consistence Hashing Ring, with quorum-like mechanisms, so you can have strong consistency and dont loss data(similar to Cassandra and Riak) and also have some low latency and high throughput using Redis or Memcached Behind. Dynomite is written in C and its a proxy, it uses the twitter twemproxy as base solution. Replication is a aymetric, dynomite has a java client called Dyno with has Token Aware load balancing. On the consistency side you can do: DC_ONE: Sync same AZ, Async other Region or DC_QUORUN: Sync to the mun of the quorum.
Performance is amazing, check this benchmarks by Netflix folks. They used a R3.Xlarge instance with replication factor set to 3 in 3 amazon zones, they used in front of Redis and did some set of GET and SET operations. The ratio between reads and writes was 80% reads and 20% writes(pretty much Netflix scenario)



Dyno Client Features

One of the great things about the client(Dyno) is that you can choose the client you want use, so for redis you can use Jedis or Redisson for instance but since dyno is modular you can code to integrated other clients if you like it more. Some key features in dyno are:

    * Connection pooling of persistent connections - this helps reduce connection churn on the Dynomite server with client connection reuse.
    * Topology aware load balancing (Token Aware) for avoiding any intermediate hops to a Dynomite coordinator node that is not the owner of the specified data.
    * Application specific local rack affinity based request routing to Dynomite nodes.
    * Application resilience by intelligently failing over to remote racks when local Dynomite rack nodes fail.
    * Application resilience against network glitches by constantly monitoring connection health and recycling unhealthy connections.
    * Capability of surgically routing traffic away from any nodes that need to be taken offline for maintenance.
    * Flexible retry policies such as exponential backoff etc
    * Insight into connection pool metrics
    * Highly configurable and pluggable connection pool components for implementing your advanced features.

Installing, Configuring and Running


Cheers,
Diego Pacheco

Saturday, September 19, 2015

NetflixOSS: Installing and Running ICE

NetflixOSS ICE is AWS COST monitoring solution. Similar to cloud watch but IMHO is better. There are some issues with ICE first of all is kinda of buggy at least right now - something you got some NPE and is hard to figure out what you did wrong.

Second - its kinda of slow to run first time it take some time because it does lots of processing to figure out the costs.  You need have your billing configure right on AWS and also you need a S3 bucket to store the billing data so ICE and read and show the charts for you. We will install ICE on Amazon Linux OS(CentOS based).

Installing and Running ICE on Amazon Linux OS

Beware of the policies - you need add some policies in S3 in order to get this working right - in aws you can get your tight policies this are the policies are generated for me for you might be different so keep that in mind. You will need to have your amazon Key and Secret as well so when we run ice you will pass that by parameter.

When you run ice you will see something like this:



Have Fun :-)

Cheers,
Diego Pacheco

NetflixOSS: Installing and Running Vector

Netflix Vector is a great monitoring tool for linux boxes.  Today i will show how to install, configure and run Netflix vector. We will do that on the AWS cloud.

I will do the installations on the Amazon Linux OS (CentOS Based) but IMHO you can use other OS if you want like Ubuntu.

Installing and Running Vector on Amazon Linux OS

First of all - You will need create a amazon instance, select amazon linux and them once you create the machine you need ssh into the box, them you are ready to install vector.

As you can see we installed PCP because vector is just a UI - IF you want you can install PCP in other boxes and them just have vector in one HOST. This is useful because them you can have a central repository for monitoring and you can see all in first place.

When you open vector in you browser will look something like this.


Have fun :-)

 Cheers,
Diego Pacheco

Saturday, September 12, 2015

NetflixOSS the DevOps Stack for Microservices


There are lots of people talking about microservices. IMHO I don`t most of people get it - Microservices are about Isolation, Idependence and Anti-Fragility and this are principles most of frameworks did not have. So the botton line is in the end of the day if you dont have this you are just doing OLD SOA ou even worst you might just be doing WebServices.

Netflix get it. All the components are build guided by core architecture principles and anti-fragility is on the heart of this components.


There are other stacks that are very promising like Akka and Twitter Stack(Finagle) but the main problem with AKKA is that the Operation part is payed and is not really close to what Netflix has to offer.  Besides that akka has another problem - The Programming model is Actors, dont get me wrong actors are great by is not a generic model for everything and does not work well with the service idea.

NetflixOSS gets operation right, is all designed to be Observable i mean in sense of Observability - like you can go there and see whats happening. There are logging, dynamic configuration, monitoring, drivers, load balancers and all sorts of mechanism your stack need it. There are so much emphases on operation not only on building thats why NetflixOSS is ready for the devops ERA because it acknowledge ops and take it into account and thats is something very different that you dont see in standard-  pre 2010 solutions for services.

The Architecture Principles

* Separation of Concerns
* Cloud Native
* Microservices
* Everything is broken and fails constantly
* De Normalized Data
* Chaos Engines
* DevOps: Run what you wrote, Anti-Fragility, Immutable Infrastructure, Failures are Opportunities to Learn, Blameless Incident Reviews
* Commitment to Continnous Improvement

The Main Architecture

This is the NetflixOSS Service architecture. You have your devices or service consumers that are connected to the internet and they talk with AWS Elastic Load Balancer and this is call the Zuul witch is just a poxy like HAproxy them will talk to services. Netflix makes difference from internal services and external services - external once they call edge services. All services are isolated and have they own databse and they dont access a central shared database.



The Core Middleware

Netflix has a stack for microservices and operations around it. The Key components are:

* Karyon -> The Nucleus of the microservices - It uses RxNetty as server
* Ribbon -> Java IPC driver to call other services - you can do rest calls with it and uses RxJava
* Eureka -> Discoverability solution Netflix built.
* Hystrix -> Resiliency, Circuit Baker, Timeouts solution -> Wrap app code that is danger with Commands.
* Turbine -> Visual Stream Aggregator for Hystrix - you can see failure and timeouts and circuit breakers at runtime
* Archaius -> Dynamic Configuration Manager for the JVM
* Governator -> Netflix uses Guice and here are the abstractions and wiring utilities.
* Zuul -> Proxy Server that does simple routing and security

For the Operations

* Asgard -> Kinda of Jenkins for the Cloud - Can create clusters, ASGs, ELBs
* Aminator -> Python solution to bake amazon AMI images
* Servo -> Monitoring solutions
* Ice -> AWS Cost Visualization and Monitoring
* SimianArmy -> Chaos Testing - Tear down data centers, instances, burn CPU
* Vector -> Monitoring JVM and servers at runtime

Next posts i will cover some of this solutions with code and examples - also will provide github code working :-) NetflixOSS is great but the documentation is not 100% and sometimes you really need debug and hack the code to understand whats going on. SpringCloud is the Netflix Stack(some very small part of it) with Spring not Guice and has some documentation.

Tech Posts with Code

* Microservices with NetflixOSS: Karyon, Ribbon and Eureka part 1 
* Microservices with NetflixOSS: Building and Running Eureka part 2
* Microservices with NetflixOSS: Karyon and Services part 3
* Microservices with NetflixOSS: Ribbon part 4

Cheers,
Diego Pacheco

Chuyên mục văn hoá giải trí của VnExpress

.

© 2017 www.blogthuthuatwin10.com

Tầng 5, Tòa nhà FPT Cầu Giấy, phố Duy Tân, Phường Dịch Vọng Hậu, Quận Cầu Giấy, Hà Nội
Email: nguyenanhtuan2401@gmail.com
Điện thoại: 0908 562 750 ext 4548; Liên hệ quảng cáo: 4567.