Unclaimed: Are you working at Hive?
Hive Reviews: 4.2/5 — Solid Choice
Hive is an all-in-one project management tool developed to “help teams move faster” regardless of how they work. Features are created based on users’ requests and are updated weekly, making Hive the world’s first democratic software platform. It’s best known for its capabilities in project management, time management, team collaboration, automation, and an array of integrations with third-party software. Hive is free to use for solo users and with premium versions available to teams and enterprises.
| Capabilities |
API
|
|---|---|
| Segment |
Small Business
Mid Market
Enterprise
|
| Deployment | Cloud / SaaS / Web-Based, Mobile Android, Mobile iPad, Mobile iPhone |
| Support | 24/7 (Live rep), Chat, Email/Help Desk, FAQs/Forum, Knowledge Base, Phone Support |
| Training | Documentation |
| Languages | English |
Compare Hive with other popular tools in the same category.
Fantastic and attractive UI Straightforward when it comes to tracking,assigning and monitoring projects Hive helps me keep track of all our projects and people connected to these projects
I don't have anything yet to state as a dislike I am satisfied in using hive
Managing the projects and getting the time logs
Hive is the best Datawarehouse and open source. Easy to use. It has a query language HiveQL. The syntax is same as SQL. So it is easy to write the hive queries and easy to get reports and business insights. Best Olfor analytics. And easily integrated with Spark, Hadoop and cloud also.
Higher latency is the drawback. Developments should be made to improve the latency of the complex queries.
Helping us to get business insights and reporting. Helping us to analyse the data and draw some conclusions to improve the business.
Best thing about Hive is it alows us to write the code in sql for processing and managing the data. It allows us to partition and distribut the data among the cluster machines. We can also choose execution engine like spark, tez,mapredice while rinning in hive. If you your processing and analysing the bulk volume of batch data then hive is best.
It supports only structure data and we can't do updates too. Only supports OLAP.
We are managing our data warehouse for analytical purpose using hive. We used to partition and cluster the tables for better parallel processing and less shuffling so the processing will be optimised. We used to read the hive tables with spark too for better computation and speed.
Good querying on hive databases.and easy to create schemas
It does not allow datatypes conversion. As it will result in lost of data
Querying on hive dbs. Also hs2 and hive metastore resulting fast results
Schema on any format HDFS files. Easy to download the data. A complete tool similar to database tool like toad.
Performance,sometime it is very difficult to run queries. Gui can be improved with more user friendly options
Data processing for regulatory reporting ...maintain lineage
I was a top fan of Impala for a while until I reached a series of limitations that were impossible to overcome. I work a lot with arrays and just the fact of being able to use array_contains in impala made me switch to Hive. Also, we are moving fast on the direction of self made Macross for hive that let us do complex queries without lateral view explodes
Session creation takes a while and speed is quite slow when comparing to Impala
Complex data analysis with tables that have several billion rows by partition
It is highly flexible in configurations. So many options to load data from- directly from linux file system or hdfs. You can create external and managed tables. One fun feature is that you can shoot bash commands from hive as well
It cannot be used for streaming data. Error logging can be improved so that error tracking and resolution can be more efficient.
It is used to transform and process Big Data datasets in batches. It can handle TBs of data. Push predicate feature has greatly improved the performance of the queries and the developer doesn't need to think about it anymore
Hive is great for handling logs in big data projects. We are using the same in our project and it is great for using joins and grouping which is very difficult and tricky in map reduce. It has a lot of udf packages and it is very easy to add new udfs. We were also using bucketing and clustering to optimize the query. Concept of external tables and the way we can manipulate data even when table is deleted from hive is really amazing. Lot of connectors available in the market for different softwares.
The thing which I dislike is latency and the way it saves data. While inserting data I have to wait a lot of few records. Compiler execution plan is very immature as it does not do proper query optimization. Though the community is working fast for overcoming quickly but I think it will take time for hive to be
We are using hive mainly for saving our logs. it helps us to keep track of what records are inserted, which records have failed and what are relationship between them. we are using tableau for analyzing data .
The progression of features, speed, etc brings me the strategic confidence I need in the SQL in hadoop space.
At this point, everything is on pint & theories it is great in hive 1.2
Deriving value from masses of unstructured & structured data.
1) ability to handle PetaByte scale of data 2) ability to work with near SQL query language 3) schema on read capibility
1) optimizer technology is still maturing
1) handling huge datasets 2) handling semi-structured and structured data sets with the same tool
The SQL-like syntax makes it easy to use
Its reliance on mapreduce makes it sometimes very slow even for a simple task
Run specific queries on hadoop files
The ability to view HDFS data in a relational format and easily query it through HiveQL
The fact that it uses MapReduce whether you query a pre existing table or a perform a complex query. Tez helps with this issue. Also the inability to delete/update data is a real issue and forces other services to be used eg HBase.
The ability to use Hive on HUE is perfect. We are building a platform for data scientists (prefer GUI to shell) to perform analysis so removing the need for command line is excellent.
It is very simple to use because you fill like you use simple SQL language for querying data. When I just started I didn't have any experience with Hive and in like one week I was able to query big data and do some analysis. In a month I was able to administrate data and create my own databases with the useful data. . .
Not so many implemented functions in the Hive. There are very useful Window functions but it's not enough. . . It's not that simple to modify data inside a table. . .
Analyze every day and every hour or even every minute user experience, user behavior in application or web client , etc . . .
Hive is one of the Apache projects. it is a data warehouse software which facilitates querying and managing large datasets residing in distributed storage. It provides a way to enable easy data extract/transform/load (ETL) Some of the nice features include (1) a simple SQL-like query language, called HiveQL, that enables users familiar with SQL to query the data. But it is a bit different from SQL standard. For example, HiveQL can also be extended with custom scalar functions (UDF's), aggregations (UDAF's), and table functions (UDTF's). (2) You can define your own read or written data format called Hive SerDe. (3) Hive can be run on Hadoop and HDFS. It has very good scalability. Personally speaking, I use hive mainly for ad hoc quires and reports. For BI reports Hive is the best since you can reuse all the SQL that you have done for traditional data warehouses. Also with Hive Server2 you get a real JDBC support so you can plug your BI tools to it. Many more SQL features like cubes, rollups, windowing, lag, lead, etc are being added to Hive through Hortonworks Stinger initiative. Hive also produces very compact code, which is always good for reading and debugging.
Too large code base. It is hard to maintain and support. And, there are too many configurations. If you take a look at the HiveConf.java, you will be confused with so many configurations there. It is easy to get lost when you configure them. And, if you configure some of them in a wrong way, you may suffer from bad query performance.
We are developing Hive. Hive is part of our product
Apache Hive is a tool built on top of Hadoop for analyzing large, unstructured data sets. Most BI and SQL developer tools can connect to Hive as easily as to any other database.
Unable to cancel a running query. Query tuning is difficult compared to RDBMS
We had a requiement to scan a large dataset for our predection algorithm. Initially we used RDBMS but the performace was very slow and user where not happy with it. We replaced RDBMS with the Hive and we are able to see a drastic improvment in the performance.
hiveql is more like SQL and really easy to learn
doesnt work good if you want a low latency queries
performance for 1TB of data
If you know SQL you will be able to get Hive really quickly. Lots of the same functionality but not exactly SQL. Easy to create tables and start writing queries allowing you to dive deeper into your data.
As with all Hadoop tools lots of knobs to tweak. Takes a good bit of time optimize and finely tune your Hive install.
Putting structure on unstructured. Once we chose hive to accomplish the aforementioned task we were able to bring our data to our data scientists quickly. An easier degree of acceptance to the Big Data idea.
Its exhilarating how well it seamlessly connects work, through its group messaging that spurs collaboration, its flexible project layouts, customer insights and reviews that help in strategy.
One thing to note is that its heavy on consumption and it requires enough ram to slow some integrated apps down. That's an issue that needs solving.
Its been using data analysis of the workflows to expose how work is actually being done, what's needed to optimize operations, how to schedule, allocate and much more. Its been a great help at getting everyone performing and moving as per timelines and objectives.
Hive offer tool's for project planning, including the ability for track project and timeline, create and assign tasks and manage resources.
I have no dislike towards hive since is the one I'm using
It allows users to analyse large amount structure and semi structure data
Hive is a very valuable tool as it provides wrappers for data analysis and querying on Big data for organizations with huge amounts of data to be processed. It is built on top of Hadoop and makes SQL query building and storing quite convenient!
The biggest disadvantage of using Hive is that it does not provide or offer real-time queries and especially for row level updates as the latency is quite high in Hive.
Hive solves the problem of big data processing and analysis for me and my company. We are able to process and analyze huge amounts of data with the help of capabilities provided by Hive. It also allows parallel processing which makes it quite fast to use.
File formats for optimizations, external
cannot analyse unstructured data, not much faster when it comes to complex operations on very huge data
adhoc analytics and batch analytics on big data
The data distribution and data processing in hive is very good in hive. DDL and DML functions in hive are better than conventional sql databases. Hive is better known for fault resistance.
The data output on hive is very slow. The data processing is very slow and the output is delayed due to this slow processing. Also the syntax is bit complex than conventional sql language.
Fault tolerance is very good feature which helps in data storing. Incase one database is lost, there isna backup created and this database can retrieved easily and intact.
the ease of use and its interface is best
running of queries in slow mode of map-reduced
storing the large data and retirving it through queries
Caso contrário, se você escrevê-lo como um SQL normal, pode levar horas para processar Mas é um pouco diferente do padrão SQL Pessoalmente falando, eu uso hive principalmente para ad hoc quires e relatórios é um software de armazenamento de dados que facilita a consulta e gerenciamento de grandes conjuntos de dados residindo em armazenamento distribuído
O ajuste de desempenho é difícil e torna-se difícil para consultas complexas, ele ainda tem alguns bugs, como todos os dados que vão para um único redutor, o que pode levar a retardar os resultados da consulta. -> Algumas das operações SQL não funcionam na colméia, como as associações de não igualdade, os dados não podem ser atualizados, mas teremos que reescrever
Estamos desenvolvendo o Hive Para pessoas que estão acostumadas a escrever consultas SQL, seria muito bom usar o Hive em cima do hadoop para arquivos armazenados no HDFS. Atividade do Dumping Site Dados de streaming de Big Data, bem como logs de dados no Hive Estamos desenvolvendo o Hive
- Easy to use interface - multiple clients (CLIs) - easy to debug issues with the help of fully descriptive logs - constantly the product is being improved to meet all the DB developer requirements - can be accessed from multiple applications - access through knox for additional security - no indexing - multiple file formats - the tez architecture
- authentication gaps - issues when routing through zookeeper - not as matured tool as the regular database tools
- BI team is helping all the enterprise users to ingest and access data from hadoop - most of the users are well versed with standard sql tools - to make hadoop enterprise wide solution we are training all users with hive
Ease to get started. Leverages sql knowledge. Has reasonable documentation. Fast to write queries.
Documentation sparse in some areas such as datetime formats. Queries run slowly and often fail to complete.
Preprocessing for machine learning pipeline. Running ad hoc queries on customer databases to generate high level summaries.
Its very user friendly. Easy to install and use. I like the interface very much .
Nothing I can think of. I had a great experience working on HIVE and was satisfied with all the features as it met all my requirements.
I was working on a school project as a part of Big Data course and executing queries with HIVE made the whole project lot more simpler.
Hive has a simple and intuitive interface and gets the job done.
So far Hive has met and exceeded all my expectations.
Working on a Hadoop system to determine recruiters that are spamming members too much.
The syntax of hive! Its almost SQL so its easy to use. External tables, partitions, buckets, UDFs all the features I like to use with hive. ORC data format occupying lesser space and retrieving the data much faster. Learning curve looks easier as it is similar to SQL but hold on! you must learn all the features of hive before writing a big hql to join multiple hundreds GBs tables and fetch results. Otherwise if you write it like a regular SQL it may take hours to process. So hive is always at its best when you set the optimization parameters before you run your scripts. Also its complex datatypes make hive more useful than other RDBMS.
Hive is comparatively slower than its competitors. Its easy to use but that comes with the cost of processing, If you are using it just for batch processing then hive is well and fine.
Generating datasets from huge files for reporting purposes.
It's performance using distributed computation
Limited options for query performance optimization
It is very good for OLAP related tasks