Showing posts with label aws. Show all posts
Showing posts with label aws. Show all posts

Thursday, January 2, 2020

Amazon Builder Library: Review Notes

During AWS RE:Invent 2019 amazon release Amazon Builder's library. Which are world-class Distributed Systems, Engineering, Cloud-Architecture, DevOps class and lessons learned through experience on how to build reliable and scalable systems. The library is really great, some articles are a bit extensive like 30min read but they are really well crafted and provide great insight on amazon best practices. Today I want to share my review notes on the Library. I highly recommend you read the library. Amazon content has links with their products and also links to further readings in the sense of theory & tools. It took me a weekend to go through all the articles(13 so far) and this time was well invested I had lots of fun and learn a lot.



Review Notes

My Amazon Builder's Library Notes

Companion Deck



I hope you enjoy!

Cheers,
Diego Pacheco

Wednesday, January 1, 2020

Amazon Aurora: Papers Review

Amazon Aurora is a very interesting piece of engineering. I want to share my review notes on 2 papers related to amazon aurora. I was wondering for a long time, 2 years if I should start publishing my reviews or not. I do structure learning and study for a long time for more than 15+ years, however, I often take notes on the papers and videos I see in my Evernote. After 2 years of thinking I decided to share my papers review notes. I believe this decision will improve my learnings because I will write even more about the things I study and read and often do in my day-by-day work. I don't work at Amazon however I like amazon a lot, they are a reference for me in regards of SOA(Service Oriented Architecture), DevOps, Cloud, and NoSQL/NewSQL databases now, which aurora they proved to me that cloud-native database it's a real thing and its not BS marketing.

Papers

Engineers often dont read papers, which is a big mistake because they loose often Deep Dive tech content related to Algorithms, Datastructures, Techniques, and lessons learned. So If you are reading this, start reading papers. The 2 amazon papers I'm talking about are:

Amazon Aurora: On Avoiding Distributed Consensus for I/Os, Commits, and Membership Changes
https://dl.acm.org/doi/10.1145/3183713.3196937

Amazon Aurora: Design Considerations for High Throughput Cloud-Native Relational Databases
https://www.allthingsdistributed.com/files/p1041-verbitski.pdf

I highly recommend you read the papers if you want to learn more about cloud-native databases.

Reviews

Amazon Aurora: On Avoiding Distributed Consensus for I/Os, Commits, and Membership Changes
https://github.com/diegopacheco/notes/blob/master/paper_reviews/aws_aurora/diegopacheco_paper_review_Amazon_Aurora_On_Avoiding_Distributed_Consensus_for_IOs_Commits_and_Membership_Changes.pdf

Amazon Aurora: Design Considerations for High Throughput Cloud-Native Relational Databases
https://github.com/diegopacheco/notes/blob/master/paper_reviews/aws_aurora/diegopacheco_paper_review_Amazon_Aurora_Design_Considerations_for_High_Throughput_CloudNative_RelationalDatabases.pdf.pdf

I hope you enjoy it. More to come.

Cheers,
Diego Pacheco

Tuesday, March 12, 2019

Running AWS ES Open Distro locally

AWS Open Distro for ES is an open source distribution from Amazon for ElasticSearch cluster and Kibana. Today I will show a simple script I made in order to run ES/Kibana locally without pain. You will be able to run the solution without worrying about any configuration. You need to keep in mind this is a developer sandbox script, this is not production ready ES cluster config. AWS ES Open Distro is enterprise-grade because it has advanced capabilities like SQL, Alerting and cluster diagnostics, I like the cluster tools a lot.

The Script


The script uses docker and link 2 containers(es and kibana). The script also creates an index in ES called Twitter having 3 documents getting indexed with the following props: user, post_data and message.  The script copy a custom kibana.yml file into kibana container in order to pass proper docker DNS / Container link to ES container - by default config looks up to localhost and I have to change to es DNS(same name used on docker link).

Checking ES and Kibana

ES: https://localhost:9200/_all 
Kibana: http://localhost:5601/app/kibana

ES (_all)



Kibana


Perf Tool

Node Analysis


Cluster Thread Analysis


Cluster Network Analysis


Cluster Overview


When you are done you can simply do:

docker kill es kibana

Cheers,
Diego Pacheco

Monday, February 11, 2019

Running Istio on EKS

Besides all network overhead, Istio offers very interesting trade-off for sacrificing latency and network overhead for developer productivity and stack independence. In previous posts, I blogged a lot about kubernetes, Istio, Aws, Kops, Eksctl and EKS. Today I will show how to run Istio in AWS using EKS. Keep in mind EKS don't support Alpha* Specs right now(v1,v2 or v3) so some demos from the istio best selection of slideware won't work. But is possible to have istio installed and booking app running.




Running Istio in AWS with EKS and eksctl



Cheers,
Diego Pacheco

Wednesday, February 6, 2019

Running k8s on EKS

EKS is the new AWS managed Service for Kubernetes launched at last Re-Invent 2018.  EKS is not available in all regions right now. EKS an option for those who don't want to use KOPS.  For this blog post, I will show how to easily set up a kubernetes cluster in AWS using EKS.  EKS has some benefits, first of all, is a managed service that you are not locked in since the API is kubernetes based so you can easily migrate to other kubernetes installation or even other kubernetes installation in other cloud vendor or on-premises.




EKS Benefits
  • Managed Service without lock-in (Kubernetes API, specs, kubectl)
  • There is no Control Plane management (Multi-AZ, Automatic patches, and updates)
  • Secure by Default (Secure and encrypted communications between worker nodes and master)
  • Conformant and Compatible (EKS runs certified kubernetes. Compatible with standard k8s envs)
IMHO this is what most of the companies are looking for, in other words, *Serverless*. 

The Pain Points

I don't want to paint a rosy picture of EKS like any new technology there are problems such as:
  • Poor Documentation - very few docs
  • Lack of Observability - It's a kind of black blocks compared with Kops.
  • When you do something wrong and har to figure it what you did it wrong.
  • Error-Prone - Is very easy to make mistakes(eksctl fix that)
  • In a Long Run is more expensive than running on EC2 with Kops
  • Alpha APIS is not supported. 
It's important to keep in mind that EKS is a very new service and it should get better and the time passes. 

Running Kubernetes in EKS using EKSCTL

EKSCTL is a great tool. Written in go, makes eks cluster creatin way less error-prone and many simples to get a cluster up and running. So let's get started! Basically, we will install the AWS Authenticator and EKSCTL them we can create the cluster and deploy nginx in kubernetes after that we can access the nginx application in the browser(before that you will need to enable the SG access for your IP). Them we destroy the cluster.


Creating a cluster














 Deploying nginx in Kubernetes












Nginx up and running on AWS(Need to enable 80 port SG access)

















Nginx AWS ELB Created by EKS and K8s with EKSCTL



 K8s Nodes Running on EC2 via EKS and EKSCTL





















Destroying the cluster







Cheers,
Diego Pacheco

Tuesday, February 5, 2019

Running Istio on AWS with Kops

In previous posts, I show how to run Istio in Minikube and with Docker-Compose/Consul in local env, today I will show how to run on AWS using KOPS.

This installation is Linux based(Ubuntu), I'm running all commands from my local-desktop, if you don't use Linux(shame on you) you can create a virtual-machine on AWS with ubuntu and run this commands there, also is possible to run Vagrant with Linux and them run this commands on Vagrant box as well. Istio runs smoothly in AWS with Kops. You don't need much, pretty much 3 machines(1 master node, 2 minions).  Keep in mind this is not a production-grade setup, for production, you should be running with 3 masters at least for High Availability.




Installing and Running Istio with Kops



Master and Worker nodes on AWS EC2 Console










Istio Metrics in Grafana

















Jaeger - Distributed Tracing






















Kiali - Observability

















BookInfo ServiceMesh (4 microservices) running on Istio / Kubernetes in AWS

Prometheus(Cloud-Native Observability) - Metrics, Dashboards, and Alerts 

















ServiceGraph























That's it - I hope you enjoyed.

Cheers,
Diego Pacheco

Running Kubernetes on AWS with KOPS

Kops is the best way to have Kubernetes running in AWS. Kops allow us to install kubernetes in EC2. Kops is written in Go. Kops helps us to create, update, maintain and destroy kubernetes clusters on aws. Kops also supports GCP(Google Could Platform). Kops has some interesting ability to generate terraform files if that's your you thing :-).  For this blog post, we will be using AWS ELB as DNS so we won't be using public DNS records which are done by setting carefully the name of the cluster - which need to end with .k8s.local. Right now is way faster to spin up a kubernetes cluster with Kops rather than EKS.



Installing and running Kubernetes in AWS with KOPS



Cheers,
Diego Pacheco

Monday, July 30, 2018

Experiences Building a Cassandra Orchestrator(CM)

Cassandra is a kick-ass NoSQL database. Battle tested by Netflix, Apple, Uber and so many great cloud scale companies. If you want to run Cassandra in production you need to buy DataStax Enterprise or engineer your own solution - since Cassandra community is not enough. My company loves open source and my customers love open as well so we decided to build our own Cassandra orchestrator for a lack of better name we called it CM(Cassandra Manager).  CM runs on AWS(Ec2) and managed single and multi-region clusters for Cassandra. CM has java interfaces so even today it just runs on AWS it could easily be ported to in containers or other fabric runtimes, which may happen on the long-term future. So it might sound crazy when you think about to build engineering around Cassandra but is actually not several companies in Silicon Valey like Netflix and outside of the valley do similar things or Databases being NoSQL or Relational. Automation is must if you want to scale, automation is way more them "automated deploys" that's easy the ultimate automation is when you automate your operation and them allow you to truly scale.

Why Build CM?

Basically, we won't be able to Scale AWS Single and Multi-Regional deploys and Operations so we don't need to code any Jenkins or bash script in order to serve more clusters. Before CM we were doing backups with Jenkins jobs and we did not have ASG on top of Cassandra nodes so cloud ops team always had issues when a node goes down and had to run bash scripts to put the node back online with CM all that changes. The single region is not that complex but when we talk about multi-region things gets more complicated. So automation is really the solution.

Who we build complex system on the fly? 

In order to Build CM we had to some POCs and also drawn some Color UML diagrams based on Responsibility-Driven Design(RDD). This might sound a bit old school and design up to the front, however, there is no way to build such a system without some up to front thinking and strong design and architecture leadership. My team works with Kanban and I do multiple roles in my team playing as Tech Manager / Agile Coach, Engineering, Architect, and DevOps Engineer Testing and Developer Support. My team does multiple roles as well so an engineer in my team does Software Engineering, DevOps Engineering, Testing and Developer Support. I and my team work with the idea of Self Service systems where a developer could easily consume what we do, so there are generic Jenkins jobs and good documentation is the place(Internal Wiki).

We use Kanban and track stories per week with simple math items/weeks we easily could keep develop and deliver CM in time. CM took 3 months to get it done. Thanks to our Kanban management, and internal system design with RDD and Color UML we were able to work normal hours. However, I underestimate the effort to get the system STABLE and the stabilization period that should be 2 weeks become 2 months. During CM tests me and my team found ~30 bugs in CM. I work with a small team right now(me and 2 people) when we start coding CM was 4 people and 2 guys left so after stabilization all gets a bit complicated and we had to work several hours in order to keep with new team and lots of CM bugs this was a bit stressful but we overcome that and deliver to production :D

It's all about Availability

When I was thinking about CM I knew we will need to make CM highly available. In order todo, we made CM optinional so Cassandra can operate without if CM dies or crashes after deploy. In case CM dies or crashes Cassandra keep operating normally and since CM is an orchestrator developer and application who w/r in cass don't see cm. So if CM goes down we lose some operations like:

  • Backups
  • Repairs
  • Node_Replaces
CM has an Autoscaling group for it and persists its internal state in S3 so we build a recovery process for CM in case that happens ASG will spin a new ec2 instance and CM will keep running again. CM was built with multiple layouts in mind so you can have 1 CM for several Cassandra clusters since CM has it own thread Pool or you can have 1 CM per Cassandra cluster, this is great because it provides more availability and allows us to save costs in low production environments like DEV and TEST where we can use shared CM or in production for small clusters we can share CM as well. For clusters that are sensible, they can have they own CM. This is so true because we had an outage in production with CM, however, Cassandra was not affected and was running fine. All this can be seen as Reliable or Anti-Fragility system and as Failure Degradation mode(meaning CM fails don't make cass fails) all great Netflix ideas we stole :D 


CM Philosophy - Self Healing and Self-Operating

CM was a set of philosophies like Self Healing and Self-Operating system. Ops teams are used to Operate systems in production due to the lack of built-in philosophy. CM has a different philosophy so basically, you don't operate CM - CM does everything by itself(Unless there is a BUG in CM - them you need RECYCLE / OPERATE CM).  CM was highly inspired by Dynomite, Dynomite-Manager, and Priam from Netflix, Although CM has a different architecture. I worked a lot with Dynomite and Dynomite manager on the last 3 years and this experience was very important and allow me to get it here.

Why not use Priam? Priam is great don't get me wrong. However, for my use case I need to support Cassandra 2, Cassandra 3 and soon Cassandra 4(when gets released). Priam has support for Cass 2 and cass 3x, however, Priam has branches and Netflix is migrating the whole fleet to latest Cassandra and we want to have a more long-term support for cass 2 and cass 3 and have our time windows. Priam is a co-processor also know as sidecar and Cassandra Manager is an orchestrator meaning it doesn't have in the same node as Cassandra so we can precisely control what runs in each not and when.

We don't have run multiple things on same time for same clusters like backups and repairs so we want to make TAKS run SERIALIZED like 1 at the time per cluster and parallel within multiple clusters in CM - this will be better covered later on this post.

Jenkins Builds



There are any Jenkins jobs - Basically, there are 2 main jobs. 1 to bake AMIS we can bake amis for Cassandra 2x, Cassandra 3x and CM. Them we have a job for lunch CM. For Devops Engineering we use Ansible in order to provision things in Amazon-Linux OS(Our Cassandras are running in Amazon-Linux latest version - we don't use ubuntu anymore).  We use Packer in order to build the AMI for AWS ec2 - We do build in 2 regions. We create LC, ASG and EC2 instances for CM. Once CM is deploying we use CM to deploy Cassandra clusters. CM has java code in order to provision all Cassandra clusters this way we do almost all work in Java, so we have less code in bash as possible and this is great because we take advantage of great troubleshooting infrastructure and tools for Java and also the JVM compiler and good IDEs for refactoring like Eclipse.

Flacky Integration Tests

Since almost all the system is built in java it was very easy to build integration tests. We had to have 2 Tests Suites. One suite for Cassandra 2.x and other for Cassandra 3.x We were able to have the same code and same integration(end-2-end-tests) for cass 2.x and 3.x. However, we had to create a configuration that in the beginning of the sets the AMI and configs for cass 2x and later one switch to cass 3.x. My goal was to use Junit 5, however, looks like more and more Junit is moving away from integration tests and remove suite features and runners(we don't want to use legacy support thanks) so we end up using JUnit 4 which was better for integration tests.

Integration tests stated great and they easily were taking 20min or more to run since we deploy real Cassandra clusters(2x and 3x) during tests and we need to wait for aws to create objects and later on in the same suites we shut down this cass clusters. We test all major capability like backups, restore, seeds manager, restore. So as expected out tests get Flacky. Flacky tests suck. The main issues with our tests(discussed in an Agile Retrospective) were:

1. Timeouts are evil - We can't keep tracking timeouts - So we decide to switch to progressive timeouts.
2. Lack of 9 Rules: In math we can veryfied everything using math so we can check if is correct. we was not doing that in the tests and also in several tasks on the code. So if you dont 1 like that create a file - next line should check if the file as proper created.

We still have lots of work to improve our automated integration tests. What really saved us for the release was combination of Chaos Testings with Exporatory tests and checklists. I don't want to paint a rosy picture about CM so I still have mixed feeling with integration tests and we lost so much time with this.

Generic Remediation for Cassandra

Since dynomite, I had build remediation systems, So which Dynomite I was able to replace AWs AMI or scale UP dynomite without data loss or downtime. This was possible because 2 things: First because dynomie-manager had cold bootstrap feature and second because I coded a remediation process which would call DM(Dynomite-Manager) health checker and check if the node is ok and not doing anything and them kill the node wait for the node finish bootstrap and move on and so on and son. I basically refactord this code(which was and is written in java) and extend it to be a generic Gemeriation so same code wotks for Dynomite and Cassandra. We have some properties we can remediatie any cassandra node without data loss or downtime, who? Same idea we have the generic generiation process going node to node but calling CM(Cassandra-Manager) health-checker and asking if CM is not doing anything with thay node them we run a node_repair in cass so data if copied from other node this is very similar to a node_replace operation when a node dies. If you want leran more about this I made a separated post in the past its here and here.

CM Architecture


CM is built primarily with Java 8 and Inside of Cassandra node we have a Heart-Beat daemon written in Python which calls CM and decides with the node need to check in in CM(new node) or recovery(meaning node_replace) in cass.

We deploy high available Cassandra clusters where we have at least 1 node per AZ begin at least 3 nodes cluster. For multi-region, we do at least 6 nodes cluster. CM replicate its internal state to other CMs meanly in other regions in order to sync multi-region check-in for seeds management. CM was build using PEM files(Cassandra-Nodes) and we are currently looking to change this code in order to use AWS Multi-Region VPC Peering feature so we can switch from EIP and public IPs to private IPs.

Let's take a closer look at CM architecture now.

CM Architecture - More Internal Details

Basically, we have REST interfaces exposing all CM core functionality - This interfaces can be accessed with any REST tool - we build ours on a tool called CMSH which was built in Go and will be cover later on in this very post.

CM has a core class called Internal State Manager which managed all Cassandra clusters and state transitions. There is a bunch of QuartzTasks which are responsible for doing backups, repairs, restore, seeds management and so on and on. Internal State is backed up to S3 and time to time CM sends internal metrics to SignalFX. We use collected to send all OS level metrics and CM calls Injection API in Sfx in order to send application metrics.

CM has an internal Tracker class which knows what every class is doing we have an internal Framework called Step Framework which keeps tracks of all sub-task level execution state and stores metadata internal. Later one this internal state in CM is persisted in disk and time to and also sent to S3. In the future, we will backup CM state in Cassandra nodes.

When CM talks to Cassandra it uses SSH in order to manipulate Cassandra filesystem and run local nodetool for backup and repairs. All heavy work happens in cass node, not in CM - Cm just does light state and coordination work. Even s3 Upload is done in cass node, not in CM.

CM Thread Model and Issues

First Thread Scheduling Implementation was using Quartz and Reflections and we realize this was open too much treads. Basically, you want to open your threads up to the front and they just re-use them like a Thread-Pool. However, with CM we have an interesting problem since we don't know how many cass clusters we would need to deploy. For the first implementation was creating a group of threads to run each cluster operation so was basically 4 threads per group. However we were running this on Quartz Threads without the need so first refactor was move some of the work that was in quartz threads like Heart-Beats, State Request, HealthChecker to REST operation only since they only interacted with cm internal state and was a very short living process. Second refactoring was changed quartz and open just one time the pool, however, this created a problem because if we have a big pool we would run things(TASKS) in parallel that we did not want like(backups and repairs) and if we put 1 thread it would be inefficient since the machine could handle more load and we would have 2 different Cassandra clusters competing with same resources. So it was clear that we need a new thread model.


CM Old Thread Model - Thread Issues

So basically the new Thread model would need to provide the following properties/requirements:
  • Within the same cluster be serial - run tasks in sequential arrival order
  • Within different clusters - run all in parallel
  • Run CM tasks in parallel - SignalFX metrics, Internode communication, Recovery.
  • Don't break current TASKS - we don't want re-write all tasks code(quartz tasks)
  • Replace Threads from Tasks and reuse-tasks
Whats the solution? Simple we use QUEUEs and Workers. So we create and destroy Queues and workers all the time(As a Cassandra clusters do check-in(created) and checkout(destroyed). In order to make this work I had to create a QueueManager and a WorkerManager, So the QueueManager is responsible to assign Tasks to Queues and the WorkerManager to allocate workers in Threads from the pool. This solutions turn out to be more efficient and provided all design requirements we want.


CM New Thread Model - Queues and Workers

Now Quartz is gone and all threads are managed by java concurrent Executors. There are 4 main thread pools. 2 Pools for Cassandra Clusters and 2 Pools for CM. 1 Pool is for Queue and 2 Pool is to schedule recurrent tasks(like backups and repairs) on the queue. Since is a queue one you consume it there is nothing more to do there.

Footnote: Back in 2010 I built a Scala Middleware using this concepts and ActiveMQ and JBoss 5 so I basically apply similar ideas on a smaller scale.

CMSH - Build a REPL in Go


If you work with Cassandra you might be familiar with cqlsh. Cmsh is like Cqlsh but for CM and for Cassandra clusters. Cmsh is written in go, it allows us to stress tests Cassandra and create a schema, insert data and remove data and also plays a role and centralized control panel for CM so we can trigger backups and repairs from cmsh and we can see logs can all cm rest operations pretty easily. CMsh is a REPL.So its much more productive them ssh each box and run alias scripts. This is also created because we can add functionality without having to rebake Cassandra or cm amis.  Cmsh is very productive because you can connect in multiple cms with cmsh and also there is tab autocomplete, reverse back search, persistent history and much more. First was coding my own REPL later one moved to iShell which was great and save me lots of time.

Whats Next

There are many cool and challenge aspects of CM that was not covered in this post like Telemetry and Self-Healing, Multi-Region deploys and Checking, TTL Eviction for Nodes, Stress Testing CM, Unit Testing, Tuning and optimizations and much more this might be covered in future posts. Now the main focus of my team is to keep improving CM and support VPC Peering for multi-region deployments maybe later on this year we might do experiences with Containers(Docker).  although I'm the tech leader and lead engineer this is a team effort and I want to thank Jackson and Tarzan for the hard working hours and fun for working together as an awesome team and tech challenged that is CM.

I hope you guys like it, take care.

cheers,
Diego Pacheco

Monday, June 18, 2018

Running Ansible with Docker

Ansible is a great provisioning tool. However, it can be painful to get some ansible scripts right. Especially if you need some stuff with bash and Ansible. Often baking time in AWS can be pretty high. So It's better you can run ansible locally. However, running ansible local could mess up with your OS. So the best thing is run ansible in Docker. Since the docker container will be ephemeral, once you finish running the container all changes will be lost. You also will benefit from running locally and being able to figure it out quickly whats wrong.  So today I want to share about some simple project I create in order to help to do that. This is called Ansible-Docker this is an ansible sandbox using Amazon Linux.

Getting Started

In order to get started, we need have docker and git installed. Next, we need to git clone ansible-docker and then bake it. Bake just need to happen 1 time. Baked might take some time depending on your internet connection. After baking you can run "run" command which will run ansible linter and then ansible on docker image.



The Ansible Project

For this ansible sandbox, we have simple and default ansible project structure. Which you can see here on the src folder.  There is main.yml which is the file it will be run by ansible. You can see this file delegates to a git role in ansible which is located in roles/git/tasks/main.yml



The Dockerfile

Now let's take a look at the Dockerfile. So here are installing Ansible and we are using Amazon Linux Latest version as our base Docker image. As you can see on the Dockerfile we call run.sh which will run ansible as soon as the container get up. You might see a different path that happens because I'm doing some volume mapping - You can check it out here.



That's it. Now we can run ansible locally with ansible-docker. Using this ideas and scripts you can speed up your development time.

Cheers,
Diego Pacheco

Sunday, April 15, 2018

Mocking and Testing AWS APIs with TestContainers and LocalStack

Cloud Computing is the default today. I do believe the future is containers and multi-cloud solutions. However today I work a lot with AWS. There are specific endpoints such as S3, AutoScaling, Route53 and other that are among the ones I use more in my day to day work. AWS API is easy to use however not no easy to test. Distributed systems tend to be hard to test. Having quick feedback is very important for engineers. There is some kind of tasks that need to interact with AWS APIs like S3 for backups for instance. However, if you need to wait to deploy in AWS to test it because is basically impossible to test it locally then we have a problem. We need to be able to do end-2-end testing however while you are coding or doing some troubleshooting is important to do things faster. There are 2 specific projects that can help us with this task. TestContainers and LocalStack. Today I will show to use LocalStack and TestContainers together with JUnit in order to do unit tests mocking S3 API. I will show how to do this using Java8. So Let's get started!

Running LocalStack locally

In order to run LocalStack locally we need to have Python and Docker installed. After you install them we can get and run LocalStack.



After you run LocalStack docker container make sure you shut down because when we run with TestContaoiners we will be able to do it over JUnit.

Setting Up Gradle Project

Now we need to set up a gradle project. Let's take a look at the build.gradle file.



Great. Now we can proceed and work with localstack and testcontainers together.

Hacking to use Latest LocalStack Image

There is a java project on TestContainers that does the Integration between LocalStack and TestContainers we will hack that project because we want to use the latest version of LocalStack. Right now the code is using an old version. So let's take a look at the code.



The changes I did above are quite simple - I just change the docker tag from 0.6 to latest and rename the name of class + enum and that's it.  Now we can move to testcontainers test code.

Testing S3 API and Running with TestConainers

Now we can focus on the unit test and mock S3. So let's go for it.



There are some important things here we need to cover. Let's start with the Annotation @Rule. This is important because the runner will use it when JUnit boot up and will boot up the Docker container needed. In regards to LocalStack as you can see I need to pass a Specific Service/Endpoint from AWS that we want to mock.

After the Rule we initialize AWS API - It's important to note here we are passing a custom endpoint otherwise we will reach the real AWS API. Then we can use S3 API normally and all commands will go to your Docker image. So we have mocked S3 Succesful. LocalStack support most of AWS endpoints and the combination with TestContainers is killer now is very easy to test AWS specific code.

The complete code is available on my GitHub here.

Cheers,
Diego Pacheco

Wednesday, April 11, 2018

Lessons Learned using AWS Lambda as Remediation System

Today I want to share some experiences I had in the last 2 years with AWS Lambda as a Serverless solution for Microservices Remediation. This was not a single effort from me only but a team effort. The lessons learned I want to share with you today was not only in sense of Serverless but also in sense of DevOps Engineering.

DevOps Engineering is the norm today. I really can't see a world without it. Code it is the lingua franca of everything around IT. Currently, there are more and more abstractions and solutions rising for DevOps Engineering, chances are we will need code fewer things as the time pass. However, as we push the boundaries of innovation we face new problems and new problems will always require better solutions.

DevOps Engineering is great however it's not a FREE LUNCH at all, like microservices, which are great too, there are lots of COST that are introduced. There is a requirement do structure and lay down teams in a different way and also there is a need for some common components, also called architecture or platform. There are COSTS related to DevOps Engineering because when you add a new component that component has standard properties that need to be fulfilled. An I talking about requirements? No. I'm talking about the Infrastructure COST with is not only money but a set of props that need to be provided in order to you have something usable and that adds value.

The Hidden COSTS of DevOps Engineering

Besides the Hidden costs, there are explicit costs like Design, Development, Testing, Troubleshooting. However, this explicit COSTS everybody knows so I'm not focusing on explicit costs today. Let's talk about HIDDEN COSTS.

This also could be read as "The Hidden Costs of Cloud Computing if you want to avoid a complete lock into your cloud vendor" or even could be read as "The Hidden Costs of Microservices if you really want to do microservices". What are this hidden costs? Well IMHO they are not hidden but IF you are not doing this everything might sound new, so let's get to the list:
  • Deploy: You need to have something that can be deployed when you commit code or with a simple PUSH in Jenkins for instance.
  • Provisioning: Everything you do need to be fully AUTOMATED. This means not only the OS part but also the infrastructure - Here we work with solutions like Ansible and Terraform.
  • STABILITY: Since this is CORE and will be used by many microservices teams, the shared component needs to be STABLE.
  • Observability: Everything you do might have issues and you might not know what really is going on, so it's crucial to have full observability which is: Telemetry(Dashboards and Alerts), Centralized Logging, Distributed Tracking, Notifications(Slack) and so on and on.
  • Operations: So what happens if you need to change a config or tweak something? If you need to CODE or Open a ticket and have problems Houston. Automated Operation is hard but really pays off in the long run since you scale and provide a proper experience for your users(developers). Automated Operations are also known as Remediation. When you have a software to react to events and don't require manual-human-intervention. 
So every time you need introduce a new shared component on infrastructure you will need provide this props above. This is great, however, it's not free. This is also true when we talk about Databases since DBs are common shared infrastructure components. There is a tradeoff between have the best Application Design and Infrastructure Cost. 

The Problem: Zombies

My current project uses NetflixOSS Stack which is great by the way.  We do Java-based Cloud-Native Microservices. We use NetflixOSS Eureka as our mid-tear Discovery and Registry system for microservices. We run the microservices in AWS. We don't use ELBs in the top of microservices since we use Eureka but we use AutoScaling Groups. Sometimes we run into strange scenarios where the EC2 Machine is Up and Running but the JVM crashed. So the ASG would not recycle instances in this case and eventually, this scenario created Zombies. Zombies created serious issues - leading to downtimes and availability issues. 

There was a need to build a simple remediation system, which could detect this "Zombies" and kill the instances, letting the ASG recycle that instances and boot up fresh and new instances. This, of course, was not the solution for all problems since sometimes you might have a bug on your microservice or a connectivity issue(missing security group rule). So this was the main rationale to build a remediation system for microservices running with NetflixOSS on EC2. 

The Solution V1

The solution was very simple first, we basically would need to call Eureka time to time and them get the list of all microservices which was UP and RUNNING and than call the microservices Health Checker. If the Microservice health checker was returning anything different from HTTP Status Code 200 or Timing out(HTTP request to health checker) this means that service was not OK - So maybe we are talking about a Zombine. So if the Health Checker returns !200 more than 3 times we should kill that EC2 Instance and let the ASG Spin a new Instance. 

This solution was coded using Python and we initially running on Jenkins. We have several issues with running this solution on Jenkins. First, of all the solution was coded without proper timeout control so the running time was varying too much and execution was overlapping for scheduling timJenkinsnkins. Second issues were the Jenkins queue. Back on the time, we did not have a dedicated Slave so we were competing with everybody else for the same queue. This design brought many pains to my team since we end up creating an outage in Jenkins. Besides that Jenkins was not and was never designed to be a RELIABLE system so we realize this solution should be never running there. 

We picked Jenkins as the running fabric for some reasons but the biggest reason was that Jenkins was for FREE in the sense that we would not have any infrastructure or DevOps Hidden COST like I described above. 

AWS Lambda: Solution V2

AWS Lambda was a FREE solution as well. Not as much free as Jenkins was but really close. We had better reliability and since Lambda had support for Python we could re-use or code and just change the running fabric almost for free. 


So the Python code was split into 2 AWS Lambdas functions. The first lambda was calling Eureka and for each microservice, a message was sent to an SNS topic. SNS has a nice property with AWS Lambda because you can spin a lambda instance for each SNS message that arrives. The second lambda was rector to have the rest of the Python code. So second lambda was responsible for calling The Microservice health checker and then if the health checker returns something different from 200, for 3x an EC2 instance, would be killed. This system had other components like AWS CloudWatch for triggering the first lambda every 1 minute. We use NetflixOSS Dynomite as tracking and record system. Slack as Developer and Operations notification channel so when something goes wrong, i.e: Unexpected Exception or an instance is killed we pop up notifications there. 

This solution worked way better than the previous solution with Jenkins. Since we still benefit from lower infrastructure COST. You might be wondering about Dynomite, this was not a problem since we have a full infrastructure set up with proper self-service generic deploy, telemetry, driver solution. 

AWS Lambda Issues / Limitations

I don't want to paint a rosy picture of Lambda. Like everything in life, there are tradeoffs and things that could be better - so this was the pain points we had:
  • There were some issues setting up the connection between AWS Lambda VPC and AWS VPC.
  • Troubleshooting in lambda still painful, CW logs time to appear, search far from ideal.
  • There is a limit on 1k concurrent executions. After 1k you get QUEUE and after 2k your requests will be dropped.
  • This is not lambda per se issue but the code gets very ugly since there are lots of ifs. Later I realize this is a problem that could be a better FIT for an FSM design. However, AWS Step Function is not quite there yet. One day It might be like Apache Camel or some old ESB but today is very limited in sense of EIP Patterns.
  • Lambda has 5 minutes to run - it takes more time you might be terminated. More limitations.
  • The execution limit is global, in other words, it's not per lambda. Its possible to set it up but is not default.
  • There is some Hidden cache - IF you re-call your lambda in a frequency lower than 5 minutes you will see some "cache" behavior depending on how you structure, your python code. 
In general, I'm happy with Lambda as runtime fabric for Remediation Solution. Although I think an FSM solution would have a better design FIT this would also require more DevOps Engineering COST like I said before. If you get curious about FSM here are some interesting solutions:
I hope this lessons learned to help you somehow and you consider more and more write automated operation solutions because of that's the way to go and evolution of DevOps Engineering in sense of Continuous Autonomous Operations. I think AWS Lambda has a interesting Feature and Serverless in general are growing a lot and reducing some of infrastructure costs for some use cases looks like a nice FIT. 

Cheers,
Diego Pacheco


Friday, December 1, 2017

Deploy & Setup a Cassandra 3.x Cluster on EC2

Cassandra is a Rock solid AP NoSQL database. For this blog post, I will share a simple recipe to deploy a Cassandra cluster on AWS/EC2.

There is the DataStax AMI or DSE you should consider for production workloads.  This recipe I'm sharing is for Amazon Linux(CentOS based) but you can do for Ubuntu or even in Docker if you want to.

Keep in mind this is for Development / Experimentation purpose. I'm not covering proper tunning for your workload, Compaction Strategy and Keyspace design here and you also should be doing this under multiples ASG 1 per AZ ideally. So Let's get started.



Cheers,
Diego Pacheco

Saturday, November 11, 2017

Running Dynomite on AWS with Docker in multi-host network Overlay

Dynomite is a kick-ass project. Basically, allow you to have strong consistency on top of NoSQL Databases. I've been using dynomite for a while in production(AWS) and I can say the core is rock solid and it just works.

Lots of developers use Windows or Mac for instance and dynomite is built in C and it's really meant for Linux(Like all good things).  So some time ago I made 2 simple projects to get started quickly with dynomite.  Basically, the project creates a simple dynomite 3 node cluster and let you run on your local machine with docker.

There are 2 projects - One to create a dynomite cluster with Redis -- The other with Facebook's RocksDB(Experimental). So you can use it on your local machine to Debug and it works just fine. So why not go 1 step further and run Dynomite in AWS using docker? There are cool benefits if you do this approach.

Running Dynomite on AWS with Docker

Dynomite works very well on AWS but also in any other cloud-vendor or Bare metal DC. Dynomite runs on Docker just fine too. Now you can choose to run on EC2, ECS, Kubernetes on EC2 or even EC2 with docker.

The Benefits

There are many advantages do run Dynomite with docker on aws.

Here are some Benefits -- The good things:
 - COST Savings: Since you can benefit from your reservation and do better resource utilization.
 - Less Latency: Running on docker allow you to easily deploy on the same box as application and reduce network roundtrips.
 - Portability: Same docker image can be used to run anywhere also from the developer machine.
 
The Cons

Like everything in life, there are pros and cons. Here are some I found:

- Networking: Docker networking can get very tricky and hard to maintain.
- Size Limitations: Default network in /24 so it's limited to 256 ips. Offcourse you can create more networks.
- More Complex: You will have docker, docker cluster(swarm), docker network(overlay) to managed so there are more moving points of failure compared with just running dynomite on EC2 for instance.

Getting Started 

Now we will install Docker, Docker Swarm, Configure a docker cluster, Create a network overlay and run dynomite in a cluster in Docker on Ec2. Phew! Long list. :-)



We will do something very silly and simple. So will deploy a 3 node cluster. This cluster won't have sharding(You can have sharding on dynomite - just dependents on seeds config - for sake of simplicity we will not do it) or cold bootstrapping or S3 backups - If you are interested in this feature you should take a look in Dynomite-Manager.

Basically, we need do the following steps in order to get this working. These are the steps:
 1.  Create EC2 instances(Let's say 2) - Later you can automate(Ansible, Boto3, Terraform, whatever)
 2.  Create Security Groups(Use the same SG for all ec2 instances) like sg_dynomite_docker.
 3. You need open ports(SG): 8101, 8102, 6379, 2377 and any others your app might need.
 4. Them we ssh to the box and install Docker
 5. Install docker Swarm - become a master - Docker will give you the command.
 6. Do ssh to the other box and they join the master swarm node.
 7. Create a Docker network with overlay - make sure it's attachable.
 8. Configure dynomite YAML files to use fixed ips on the docker network overlay.
 9. Do docker run and run docker dynomite container 2x in 1 host
10. Do docker run and run docker dynomite on another host. That's it.

You can use my dynomite-docker project as a starting point and make the changes there because there are configs and Dockerfile done you just need change the IPs and remove the volume mapping and make sure you create the docker network with overlay as that's it. Here there is a https://gist.github.com/diegopacheco/6c75a445337e1ac29fd9ae07a16e2500 sample snipper that might help you.

Cheers,
Diego Pacheco

Wednesday, May 18, 2016

Having fun with Boto 3

Boto3 is a kick ass AWS client for python. If you are working as a DevOps Engineer with Cloud that's a library you must check it out. Boto3 improved a lot the design in comparison with boto2.

For this blog post, I will show how to do 2 simple things but yet very powerful things. Besides boto, I will show how to use Paramiko another cool library in Python which can manipulate SSH and connect to boxes and perform commands.

You will need to have an AWS account and you SECRETS in hand in order to make the code work. You can use python 2 or python 3 both can work with boto3.

Installing Boto3

For Python 2:

$ sudo pip install boto3

For Python 3:

$ sudo pip3 install boto3

List all EC2 Instances

These keys are set explicitly just for a sake of simplicity you should omit they and run an aws-configure with the AWS-CLI. AWS-CLI uses boto under the hood. It's possible to Filter by TAGS as well.

Connet to a BOX via SSH and run a Command

Paramiko needs to have the full path for the PEM file you use to create the box. Here I'm just running the Linux command hostname to get the name of the box. However, we could run complex bash script or anything you like it. We could combine both scripts to connect to all boxes.

This is very useful because easily we can create a central checker that can check for anything in all running instances we have in an EC2 Region. Boto3 has API for killing and creating instances too, so we could kill a instance if does not a match a specific requirement of Env settings.

Cheers,
Diego Pacheco


Tuesday, December 29, 2015

Building Infrastructure with Terraform

Terraform is similar to Amazon Cloud Formation. It allows you to automate the creation of your infrastructure like VPCs, ELBS, ASGs, Instances. Terraform is generic, it works with other provides besides AWS like containers and even bare-metal servers. Terraform build infrastructure, but also launch it as well.

Terraform enables infrastructure as a Code because you can describe your whole infrastructure with a simple set of declarative files. Terraform keeps track of state, so will tell you if can do something or can't do some operation. This is very cool for operation point of view. To use Terraform you just need to download the binaries for your OS and them put Terraform in your PATH. Once you have the config files you can do $ terraform apply and the magic will happen :-)

Today i will show how to build a very simple infrastructure on AWS, you just need have your credential(ID and Secret).  Main.tf is your config for the infrastructure. Variables are var you can use into your main and outputs is what terraform will output for you when its done. Download all this 3 scripts put into terraform folder, drop you pem file as well and just run $ terraform apply.
main.tf


outputs.tf

variables.tf

Cheers,
Diego Pacheco

Tuesday, December 15, 2015

Having some fun with AWS Route53

Amazon Web Services Route53 is a high available and scalable DNS solution. It can be used to route request inside and outside amazon web services as well. Route53 allow you to configure DNS health checker to route traffic to a healthy endpoint.

Route53 has plenty of features, for me best ones are:

Traffic Flow: Route end user to best endpoints.
Latency Based Routing: Routes users to region with lowest latency.
GEO DNS: Route users based on they IP location.
Private DNS for Amazon VPC:  Manages custom domains internaly.
DNS Failover: Avoid outages routing users to different locations.

For this post i will show how to install and do some basic operations with  CLI53 a command line interface for Amazon Route53.  For this post we will use CentOS / Amazon Linux Distribution.


Cheers,
Diego Pacheco

Saturday, December 12, 2015

AMI Backing: Using Ansible to Provision with Packer

Immutable Infrastructure in on the heart of DevOps. Today Backing is a very popular idea. AWS allow you to create AMI(AMI is amazon specific image is kinda of a ISO on the OS) this image can be backed in a automated fashion. To perform this task we can use Packer but Packer will take care on the AMI and publishing to AWS but you still need to install things on this linux image lets say Ubuntu or CentOS for instance. For provision you can use Ansible.

In this post i will show how to install Ansible, Packer and build an AMI image with Packer and have the full provisioning using Ansible. 



Installing Ansible on AWS

Installing Packer


Baking AMI with Packer and Provision with Ansible

So first you need define the packer configuration. You will need create a file, let's called config.json. In this file you will provide custom variable you want yo put on the backed image, whats is the target you are building, for this case will be AMI but you can use packer to build other things like Docker containers. For last you need have the provision configuration. This config will make Packer download and install Ansible and then provision with Ansible, as you can see I'm calling a ansible playbook called Apache.yml to install Apache inside this AMI.  You also need provide a base AMI ID, you can use any AMI ID i used a very basic one with just linux on it but you can use this to create more complex and layered solutions.

Now let's check the Ansible Apache Playbook.

For last we need call Packer with all this configs, so did a simple bash script that will do the work, then you can just run $ sudo ./packer-build.sh and that's it. Have Fun. IF you do everything right you will get a AMI ID in the end of the build, then you can go to the AWS console and launch a EC2 instance based on this AMI ID and will work :-)
Cheers,
Diego Pacheco

Friday, November 13, 2015

Running Ansible on AWS

Ansible is Configuration Management and automation engine.  Its has a very minimal setup, its just depends on Python - witch is good because pretty much all linux distributions came with python.

You can run Ansible with just SSH, yes with-out agents. Like other engines Ansible is based on Recipes, in Ansible world recipes are called palybooks.

Playbooks are written in a kind of DSL inside a YAML file, often called main.yml. You can have multiple YAML files and reuse a lot your provision playbooks.

Ansible runs easily pretty much in all environments BUT windows.  For windows users its easy to use it with Vagrant.  It`s possible to have pretty complex provision scripts if needed, there are lots of modules like git, file, get_url and so one and on. Ansible provide variables, so you can bind custom values on conf files also called templates.  For this post i will how to do a very simple and straight forward setup for ansible on AWS, this will be a pretty basic setup but will be great for use as a developer doing any kinda of provisioning task, so you gonna have a enviroment to test your playbooks.

You need create a AMI on EC2, for this sample i will use Amazon Linux | CentOS based. You can do with other linux distribution if you like so, once you create the box and log into ssh, you just need run the following steps.

Cheers,
Diego Pacheco

Chuyên mục văn hoá giải trí của VnExpress

.

© 2017 www.blogthuthuatwin10.com

Tầng 5, Tòa nhà FPT Cầu Giấy, phố Duy Tân, Phường Dịch Vọng Hậu, Quận Cầu Giấy, Hà Nội
Email: nguyenanhtuan2401@gmail.com
Điện thoại: 0908 562 750 ext 4548; Liên hệ quảng cáo: 4567.