Showing posts with label big data. Show all posts
Showing posts with label big data. Show all posts

Thursday, April 9, 2015

Great Data Science Center Launch at UMass Amherst Today with Photos

Today was a jampacked, fabulous day (although we did wake up to a blanket of snow and ice and it is April 9)!

Today was the Data Science Center launch at UMass Amherst with a terrific program, many friends in attendance, and great talks and panelists.

The venue was the 6th floor of the new Life Sciences Lab building at UMasss Amherst and there were refreshments upon our arrival.

The first one I saw was Professor Andrew McCallum of the Computer Science Department, who is the Director of the new Data Science Center (and we hosted Andre in our UMass Amherst INFORMS Speaker Series a while back).
He was a terrific master of ceremonies and the visionary and leader behind this endeavor as our Provost, Dr. Katherine Newman emphasized in her remarks. I snapped the photo of them hugging below.
I got to chat with our Vice Chancellor for Research and Engagement Dr. Mike Malone and Professor David Jensen of Computer Science, who also spoke in our series.

It was great to see my colleagues from Civil and Environmental Engineering, Professors Eleni Chritofa and Eric Gonzales, fellow transportation researchers.

Our Chancellor, Dr. Kumble Subbaswamy and Provost Dr. Katherine Newman had opening remarks as did Professor McCallum.
The keynote was given by Distinguished University Professor Jim Kurose, who is now serving in the CISE Directorate at NSF. He remarked how warm and lovely it was in DC with the cherry blossoms in bloom and then he landed in Bradley Airport and drove to Amherst through a hailstorm. He, as always, had some good jokes including his first presentation as an NSF program officer during which he did not thank the most important person (but thanked others).

I left the event for 2 hours since one of my doctoral students, Sara Saberi, was defending her doctoral dissertation proposal (she passed, so congratulations are in order) and the photo below is of her with the dissertation committee: Professor Adams Steven, Tilman Wolf, and Michael Zink.


 I then came back for the poster session with my doctoral student Shivani Shukla and got to see more colleagues from the College of Engineering and the Isenberg School as well as meeting a professor from Smith College and talking with more Comp Sci professors. Another highpoint was seeing Chris Hill of MIT who came up to me to reminisce about when we were at the U. of Illinois' Supercomputing Center one summer for several weeks - this was very special and I have fond memories of (massively) parallel computing, which I loved to do! Once a programmer, always a programmer!

As Professor Kurose said in his remarks, Congratulations to the Data Science Center,, and on a great launch event.

Sunday, January 5, 2014

FireEye Acquires Mandiant and Why Cybersecurity Matters

My most recent post on this blog was on networks in mergers and acquisitions and shortly thereafter there was a very interesting acquisition announcement in the cybersecurity space: FireEye acquired Mandiant Corp. This is considered one of the biggest recent security deals as reported by The New York Times, which noted that The combination of the two companies — one that detects attacks in a novel way (FireEye), another that responds to attacks (Mandiant) — comes as corporate America has become wary of relying on the federal government to monitor the Internet and warn of incoming attacks. And according to The Wall Street Journal: Mandiant and FireEye market themselves to businesses, not consumers, and focus on blocking highly skilled hackers who can evade traditional antivirus software. But they have unique specialties. Mandiant has become famous for its investigators that act like a cyber-SWAT team for companies that have been hacked. They focus on figuring out how hackers got in and removing them from corporate systems.

Mandiant is named after its founder, Kevin Mandia, who also served as the company's CEO prior to its acquisition. According to the company's website: Mandia was profiled on the cover of Fortune magazine and recognized by Foreign Policy magazine as one of the 100 leading global thinkers of 2013. The New York Times a few months ago had major coverage of the expertise of Mandiant which was fascinating.

Mr. Mandia went to Lafayette College, in Easton, PA, which is also my husband's undergraduate alma mater. From  an article on Lafayette College's website, I learned that Kevin Mandia was a computer science graduate, and also holds a Master’s in forensic science from George Washington University. He is co-author of Incident Response: Investigating Computer Crime and articles for The International Journal of Cyber Crime.

I became interested in cybersecurity and cyber crime in my research on modeling the Internet and was involved in an Advanced Cyber Security  Center (ACSC) project, Prime the Pump, entitled: Cybersecurity Risk Analysis and Investment Optimization  with colleagues in Operations & Information Management and in Finance at the Isenberg School of Management, and in Electrical and Computer Engineering, at UMass Amherst, along with one of the university's chief information officers.

I presented some of our funded research at the INFORMS Annual Meeting in Minneapolis last October and posted some information prior to the conference. Our session, organized by Professor Alla Kammerdiner,  was entitled: Big Data Analytics for Cybersecurity, and it was videotaped. INFORMS recently posted the video of my presentation: Network Economics of Cyber Crime with Applications to Financial Service Organizations  on INFORMS' great youtube channel and it can be accessed directly below.


We hope to extend this and related work through the auspices of ACSC.

Also,  Alla's presentation was on Network Inference for Monitoring Cyber-physical Systems (CPS) and it can be viewed below.
Alla ran the Boston Marathon last April 15 and is an elite runner. She heard about the bombings after she completed the marathon route and while on the Green line traveling back to her hotel.

Saturday, January 19, 2013

Are You Up to the Challenge? Big Data and SBP 2013 Conference in DC

For the past several years I was on the SBP (Social Computing, Behavioral-Cultural Modeling and Prediction) Conference Committee and enjoyed working with a great group of colleagues with the grand finale being the conference itself. I had been responsible for the tutorials  and also helped to identify several keynoters over the years. Together, with Dr. Patrick Qiang, I also gave a tutorial. at SBP 2010.

This year, since I am on sabbatical and have been and will be out of the US a lot, I stepped down from the committee.

The SBP 2013 Conference will take place in Washington DC, April 2-5, and the keynoters are wonderful: Dr. Bernardo Huberman of HP, Dr. Michelle Gelfand of the U. of Maryland, and Dr. Myron Gutman of NSF. The tutorials should also be great.

But what really caught my interest, and the news is starting to circulate via various e-lists, is the Challenge Problem. with assistance from none other than Dr. Alex (Sandy) Pentland of MIT, who was one of the tutorial givers that I had invited for the  SBP 2011 Conference. He also was one of the plenary speakers at the Northeast Regional INFORMS Conference at UMass Amherst that I was involved in (and, I do admit, it was wonderful). Dr. Pentland had also spoken in our UMass Amherst INFORMS Speakers Series.

The deadline for the challenge problem is January 31, 2013, so time is tight!

Challenge Problem

Internationall Conference on Social Computing, Behavioral-Cultural Modeling, & Prediction

SBP is offering a challenge problem for the second time in 2013. We organize this challenge to encourage researchers to combine the power of big data and the power of systems thinking in the field of social computing, behavioral-cultural modeling, and prediction. These datasets are brought to you by the MIT Human Dynamics Laboratory, with special thanks to professor Alex (Sandy) Pentland and his research team.

Problem Description

Cell phones afford a convenient platform to advance the understanding of social dynamics and influence, because of their pervasiveness, sensing capabilities, and computational power. Many applications have emerged in recent years in mobile health, mobile banking, location based services, media democracy, and social movements. With these new capabilities, we can potentially identify exact points and times of infection for diseases, determine who most influences us to gain weight or become healthier, know exactly how information flows among employees and how productivity is affected in our work spaces, and understand how rumors spread.

There remain, however, significant challenges to making mobile phones the essential tool for conducting social science research and also support mobile commerce with a solid social science foundation. Perhaps the greatest challenge is the lack of data in the public domain. There is a need for data large and extensive enough to capture the disparate facets of human behavior and interactions. Another major challenge lies in the interdisciplinary nature of conducting social science research with mobile phones. Software engineers need to work collaboratively alongside social scientists and data miners in various fields.
In an attempt to address these challenges, we have worked with the MIT Human Dynamics laboratory to release several mobile data sets in "Reality Commons" that contain the dynamics of several communities of about 100 people each. We invite researchers:
  • To propose and submit their own applications of the data to demonstrate the scientific and business values of these data sets,
  • To suggest how to meaningfully extend these experiments to larger populations, or
  • To develop the math that fits agent-based models or systems dynamics models to larger populations. The problem itself will be open-ended and encourage approaches from different disciplines, encompassing a range of applications using this data, including:
  • Social network analysis
  • Data visualization
  • Simulation studies
  • Predictive modeling
  • Qualitative studies to supplement existing quantitative work
  • Creative new applications of the data

The DataSet

Data center dynamics: The data contain the performance, behavior, and interpersonal interactions of participating employees at a Chicago-area data server configuration firm for one month. It is the first data set to contain the performance and dynamics of a real-world organization with a temporal resolution of a few seconds. The sensor data were collected by Daniel Olguin, Ben Waber, Tamie Kim, and Alex Pentland in 2007 using Sociometric Badges.
Social Evolution in an undergraduate dormitory: The data contain surveys and sensor data about the diffusion of political opinions, diet, exercise, obesity, eating habits, epidemiological contagion, depression and stress, and political opinions from 70 residents of an undergraduate dormitory. These residents represent 80% of the total population.
Friends and Family dataset: The Friends and Family experiment was designed to study (a) how people make decisions, with emphasis on the social aspects involved, and (b) how we can empower people to make better decisions using personal and social tools. The subjects were members of a young-family residential living community adjacent to a major research university in North America. All members of the community are couples, and at least one person in each family is affiliated with the university. The community is composed of over 400 residents, approximately half whom have children. The sensor data in this data set were collected using the funf open-source sensing platform for Android phones.
Reality Mining dataset: Data contain the dynamics of 75 students/faculty in the MIT Media Laboratory, and 25 incoming students at the MIT Sloan business school adjacent to the Media Laboratory. The Reality Mining experiment conducted in 2004 was the first to study community dynamics by tracking a sufficient amount people with their personal mobile phones and resulted in one of the most complete mobile data sets with rich personal behavior and interpersonal interactions. Prior to this experiment, cell phones were not powerful enough to track people.

Data access

Data is accessible from realitycommons.media.mit.edu
Users will be asked to fill out a short form and agree to privacy and data use restrictions.

Submission and Evaluation

Submissions Submissions will be evaluated based on theoretical grounding as well as use of evidence. Winners will be selected by an interdisciplinary committee of researchers and will be recognized at the conference with 1st, 2nd, and 3rd prizes. The winners will introduce the idea briefly at the conference to the audience and/or give a quick demo. Challenge organizers intend to organize winning entries into a special issue of an appropriate journal.
The submissions (6 pages) should be formatted according to the Springer-Verlag LNCS/LNAI guidelines. Sample LaTeX2e and WORD files are available from http://www.springer.com/computer/lncs?SGWID=0-164-6-793341-0.. Submissions for the challenge can be made here.

Important Dates

  • Submission deadline: January 31, 2013 (23:59 PST)
  • Notification of Winners: March 2, 2013

Challenge Problem Co-chairs

Nitin Agarwal, University of Arkansas, nxagarwal@ualr.edu
Wen Dong, MIT Media Lab, wdong@media.mit.edu

Prior work using this data

A bibtex file of references to prior publications using this data is available for dowload (righ-click to 'save as'): sbp2013challenge.bib
BibTeX files can be read by most bibliographic software packages well as LaTeX. Many free converters are available.

2012 challenge problem winners

Congratulations to winners of the first SBP challenge problem:

  • Matthew Lease, School of Information, University of Texas at Austin "Discovering and Navigating Memes in Social Media"
  • Masoud Makrehchi, Research Scientist, Thomson Reuters, Toronto, Canada "Conflict Thermometer: Predicting Social Conflicts by Analyzing Language Gap in Polarized Social Media."

Friday, September 21, 2012

What We Learned from a Big Data Decathlete and Isenberg Alum

Today, the UMass Amherst INFORMS (Institute for Operations Research and the Management Sciences) Student Chapter had the pleasure of hosting Dr. Davit Khachatryan, of  PricewaterhouseCoopers (PWC) from Arlington, Virginia.

He received a PhD from UMass Amherst with a concentration in Management Science at the Isenberg School of Management just 2 years ago and has worked on some fascinating projects.


His presentation today at the Isenberg School, "Show Me the Data: My Experience as a Statistical Consultant," captivated the audience, which included undergrads, MBAs, and graduate students from both the College of Engineering and the Isenberg School.


Big data is a very hot topic these days not only because of the volumes that are being captured from RFID technology, ATMs, cell phones, social media, etc., but also the cost of storing data has really dropped from only a few years ago, so huge amounts of data are now available for analysis, if one can make sense of the data.


He talked about the cost of working with real data is that real data is dirty and that a large part of the time devoted to a project typically entails cleaning up and organizing the client's data so that one can then get to the interesting work of modeling and predictive analytics.

Clients now want reusable, robust, and automatic solutions, rather than a 1 time solution. He spoke a lot about using SAS and SQL and the importance of computer programming.

In addition, he spoke about testing models that clients use and the existence of  "model governance" boards in corporations that check the models.

The applications that he discussed (anonymized, for obvious reasons) included two in healthcare and one in finance -- catching whether rules that are in place work for "anti-money laundering" for a bank. Interesting, when he programmed the rules and ran the model over a long time horizon's worth of data the results differed from what the bank had caught in terms of such transactions. I loved the idea of forensics in this application domain.
Dr. Khachatryan emphasized the importance of managing relationships with clients, the importance of communication and writing skills, and also that a lot of work that needs to be done is not necessarily easy or fun.

His one hour presentation began at 2PM with a nice reception (and lunch) preceding it. I left at about 3:45 and the discussions were still going strong.

It was a terrific educational experience hearing about what it takes to be a data decathlete (knowing multiple regression, understanding time series, knowing about neural networks, being adept at computer programming, listening carefully to clients' needs,  building bridges with clients, and being a good team player).


Tuesday, September 4, 2012

Big Data and Statistical Consulting -- Looking Foward to PhD Alum's Presentation!

************ Updated information -- Dr. Khachatryan will be speaking on Friday, September 21, 2012, at 2PM at the Isenberg School of Management in Room 106. A reception will precede his talk,
beginning at 1:30PM. The poster below reflects this new information. ********************

This is the time of the year that we not only gear up for the new academic year but we also start to organize talks by outside speakers

Although I am on sabbatical this year (and was also on sabbatical the 2005-2006 year at Harvard), this does not stop me from helping the students in the UMass Amherst INFORMS Student Chapter to host speakers.

So, do mark your calendars!

On Friday, September 21, 2012, just a few days before I fly back to Sweden, we will be hosting one of our Isenberg School PhD alums who will be speaking on a very "hot" topic and will also be sharing his experiences in the consulting industry.

I am delighted that Dr. Davit Khachatryan of PriceWaterhouseCoopers will be our first speaker. Dr. Khachatryan received his PhD in 2010 from the University of Massachusetts Amherst and his concentration was Management Science.

The students have prepared the poster above
  
Title of Presentation: Show Me the Data: My Experience as a Statistical Consultant 
 
Abstract
Today data are collected in wide variety of shapes and forms, including image data from surveillance cameras, click-stream and sentiment data from social media websites, financial transaction data from ATMs, in addition to the traditional ‘designed’ data from sources such as surveys and censuses. In vast majority of cases data collection is round the clock and fully automated. In addition, the frequency of data collection can reach milliseconds and capacities of data warehouses are beginning to be measured in petabytes. The challenge emerging from the repository clouds calls for dexterity in extracting the most penetrating and pertinent information from this tsunami of data. In particular, business enterprises are in need of robust, reusable and fully automated analytic solutions able to efficiently sift through avalanche of data and produce answers that will leave their mark on the bottom line. In this talk I will present my experiences in statistical consulting arena, stressing on what I have witnessed and learned when practicing the profession during the last two years. In addition, I will discuss certain challenges and paradigm shifts that I envision to be facing statistical consultants in the future.

For those of you who would like to get more information on Tips on Organizing a Successful Speaker Series, you can read my post here.


Tuesday, May 1, 2012

Simons Foundation Funds Major Computing Institute at UC Berkeley

I was delighted to read in The New York Times that the University of California Berkeley has been selected as the home of a new Computing Institute to be headed by Richard Karp with support from  Christos Papadimitrou. Papadimitrou, you may recall, is credited for the term "price of anarchy,"  which I have written about in some of my work, including in  the Fragile Networks book co-authored with Qiang "Patrick" Qiang.

The Institute is being funded by the Simons Foundation.

The Simons Foundation was established by James Simons, who had been a faculty member at SUNY Stony Brook in Math (I had an offer from the department when I was on the job market finishing up my PhD at Brown University) and then went on to establish and lead the very successful hedge fund, Renaissance Technologies. He is now one of the wealthiest people in the world but doesn't just sit on his gold.

I have written about the good work that his foundation has been doing in supporting research and education on this blog.

It is thrilling to see that this institute will be focusing on computing and will support research and applications in algorithmics and computational thinking that are now permeating so many disciplines.

I wish UC Berkeley lots of success in bringing disciplines together and in making discoveries through analyzing big data. I appreciated also seeing massively parallel architectures noted in the Times article and seeing Jeannette Wing quoted.  Yes, there are females that love to compute and to solve problems.

Friday, March 30, 2012

Thrilling New Big Data R&D Initiative from NSF-NIH

Yesterday was an exciting day for big data with the U.S. federal government announcing a major new research initiative on big data at a level of funding of $200 million.

Even The New York Times yesterday featured an article on the joint program between the U.S. National Science Foundation (NSF) and the National Institutes of Health (NIH). According to the article: Big data refers to the rising flood of digital data from many sources, including the Web, biological and industrial sensors, video, e-mail and social network communications. The emerging opportunity arises from combining these diverse data sources with improving computing tools to pinpoint profit-making opportunities, make scientific discoveries and predict crime waves, for example.

There was a great quote in the article by Dr. Farman Jahanian of the National Science Foundation’s Computer and Information Science and Engineering (CISE) directorate: "Data, in my view, is a transformative new currency for science, engineering, education, commerce and government." “Foundational research in data management and data analytics promise breakthrough discoveries and innovations across all disciplines.”


Just a few minutes before logging off for the night, to settle down and read The Times, I received the following email from Dr. Jahanian announcing the great news:

Dear Computer and Information Science and Engineering (CISE) Community,

This afternoon at a White House event, the Administration unveiled a Big Data Research and Development Initiative, which creates enormous opportunities for extracting knowledge and insights from large and complex collections of digital data. The CISE community is well poised to become an active participant in this new initiative.

NSF Director, Dr. Subra Suresh, joined other federal science agency leaders to discuss cross-agency plans and announce new research efforts to address big data. NSF will direct its current efforts to develop new methods to derive knowledge from data; construct new infrastructure to manage, curate and serve data to communities; and forge new approaches for associated education and training.

The cornerstone of the announcements includes a joint NSF-NIH solicitation on foundational research for big data. The "Core Techniques and Technologies for Advancing Big Data Science & Engineering," or "Big Data" (http://www.nsf.gov/funding/pgm_summ.jsp?pims_id=504767>) program aims to advance the core scientific and technological means of managing, analyzing, visualizing and extracting information from large, diverse, distributed, and heterogeneous data sets in order to accelerate progress in science and engineering research. Specifically, it will fund research to develop and evaluate new algorithms, technologies, and tools for improved data management, data analytics, and e-science collaboration environments.

Other announcements included anticipated cross-disciplinary efforts such as an Ideas Lab to explore ways to use big data to enhance teaching and learning effectiveness, and the use of NSF’s Integrative Graduate Education and Research Traineeship, or IGERT, mechanism to educate and train researchers in data enabled science and engineering.

For more information, please see the NSF press release (http://www.nsf.gov/news/news_summ.jsp?cntn_id=123607&org=NSF&from=news)...). We look forward to your participation.

Best,

Farnam

Farnam Jahanian
Assistant Director for CISE
National Science Foundation

Tuesday, February 21, 2012

Networks and Operations Research, Patterns and Big Data -- So How Does Money and Traffic Flow?


Add Image

Add Image

We've heard about the new geography -- now we are hearing about the new mathematics and at the prestigious Goldman Sachs technology conference that took place recently in San Francisco.

There, as noted by Quentin Hardy, writing for The New York Times, Steve Mills, IBM's senior vice president for software and systems was quoted as saying that when it comes to algorithms, "If I can do a power grid, I can do water supply..." Even traffic, which like water and electricity has value when it flows effectively. Moreover, Mills is quoted as saying in terms of cross-pollination and finding commonalities and patterns with the help of big data and algorithms that we are now: "leveraging the cost structure of new mathematics."

When I read about flows and especially in the context of transportation, electric power grids, natural resources, and finance, of course, networks and operations research and even economics immediately come to my mind since I have been researching and writing about the topic for many years.

I have always been fascinated by the commonality among problems and find that networks and the associated methodologies, such as optimization theory, game theory, variational inequality theory, and projected dynamical systems, provide a captivating medium for visualization, analysis, and computation and a way of bridging disciplines.

In our research and publications, we were able to answer several open questions, raised over a half a century ago:

1. How does money flow -- does it flow like water or electricity (raised by Cohen in 1952).

2. How are electric power generations and distribution networks like transportation networks (raised by Beckmann, McGuire, and Winsten in their classic 1956 book, Studies in the Economics of Transportation).

You can see the answers to the above questions in our papers in Computational Management Science and Naval Research Logistics, respectively.

In addition, we established, through a supernetwork formalism, how supply chain network equilibrium problems could be reformulated and solved as transportation network equilibrium problems -- thereby, constructing a common conceptual, modeling, and algorithmic framework for these two important classes of problems. That paper, On the Relationship Between Supply Chain and Transportation Network Equilibria: A Supernetwork Equivalence with Computations, was published in Transportation Research E 42: (2006) pp 293-316.

I also have a synthesis of the above equivalences in my Supply Chain Network Economic: Dynamics of Prices, Flows, and Profits book, which was published in 2006 and written that glorious year when I was a Science Fellow at the Radcliffe Institute for Advanced Study at Harvard University.

For those of you who are interested in the history of networks, with a focus on financial networks, and which includes a reference to Cohen above, please see my presentation on Financial Networks delivered at the Workshop on Measuring Systemic Risk, courtesy of the Federal Reserve Bank of Chicago and the University of Chicago in December 2010.

What seems like new mathematics (as was the case also with the new geography) to some is that now a wider circle is coming to realize the importance of research that has been going on for quite a long time, with major innovations over decades (frankly, centuries). Yes, I think it is really cool when an algorithm I implemented for predicting urban traffic flows can also be used to predict product flows on supply chain networks or financial transactions. Even major corporations are starting to notice.

Sunday, February 12, 2012

Operations Research, Analytics, and Big Data -- These Are Exciting Times for Math and Computing

For all the geeks out there who love math and computing and know the power of both to solve important problems -- your time has come and you are truly cool now.

We are now in the era of big data in which knowing and applying the right analytical tools from operations research to statistical data mining can transform decision-making.

The applications are immense and range from business to social sciences to healthcare, political science, and even policy analysis.

Steve Lohr has a terrific article in The New York Times, The Age of Big Data, which states:

If you can see patterns and make sense of the explosion of data, you are the future.

In addition, according to the article, in order to exploit the data flood, the US will need many analysts with a deep knowledge and appreciation for numbers, data, and how to extract and apply the information that is now available from web traffic to GPS data to sensor data. A report last year by the McKinsey Global Institute, the research arm of the consulting firm, projected that the United States needs 140,000 to 190,000 more workers with “deep analytical” expertise and 1.5 million more data-literate managers, whether retrained or hired.

To really know how to make use of the data for better decision-making one needs mathematical models and algorithms, and, hence, education in operations research and management science (OR/MS) is now more important than ever.

Pick your favorite area of application and with an OR/MS education and skill set you can make a difference (and have so much data at your fingertips to work with).