August 21, 2011

A warm feedback from Sigcomm

The SIGCOMM conference just finished two days ago. Papers, slides, and the video of the talks are online for free. As could be expected, there is no comparison to my experience at ICC. Despite video recording prevented presenters to move on the stage, the talks were excellent: long enough, well prepared, and in a perfect english. For every talk, many questions immediately raised and people actually debate during the coffee break and social events. In brief, Sigcomm is a conference that is worth the price (registration and travel). A series of remarks below:
  • a Sigcomm paper should present "novel results firmly substantiated by experimentation, simulation or analysis." My understanding is that "substantiating ideas" now prevails, and that the novelty has become debatable. Some ideas, which are remarkably substantiated, do not open enough perspectives. For example, deploying wireless antenna on top of data-center racks is a cool idea, but I would not include it in my list of major scientific breakthroughs. Sigcomm program committees are expected to prefer papers that are "exciting but flawed" to the "correct but boring" ones, yet exciting is not always synonyms of inspiring. In this vein, the program includes three papers related to bit-torrent. Come on, we are in 2011! How many scientists are still interested in such an overwhelmingly addressed research area?
  • Europe is back, with six papers. I already mentioned that EU-funded FP7 STREP projects match the characteristics of a competitive Sigcomm paper. This year's program demonstrate the benefits of writing Sigcomm-compatible FP7 project deliverables as all accepted European papers are (sometimes partially) funded by FP7 framework. Such fundings give the opportunity to evaluate a well-identified idea though large-scale deployment. The twenty-six other papers come from prestigious american institutions, which are probably the only places that combine a unique skill in the Art of Writing academic papers and the capacity to substantiate any idea with a bunch of outstanding experimentations.
  • I am not really into measurements, and I will undoubtedly not be. That's probably why I struggle to identify the scientific point behind the six papers that deal with measurement in the program. Indeed, it seems that the main contribution is the result of the measure, not the way these measures have been obtained. They do not present a novel super-approach to make a brand new measurement set. Rather, the idea is that these measures provide key insights of the behavior of a particular application. I agree, but does it deserve a 14-pages LaTeX-written paper? Measurement papers would probably better fit with an infographics (like this one), wouldn't they?
  • I enjoyed some presentations, especially the controversial model that explains the evolution of protocol adoptions, the scheduling of network flows in data-centers, the synchronization of multiple distant data-centers, and the reduction of redundant data transfer.


    July 8, 2011

    Leveraging collaborative projects to produce better academic research

    Opposing industrial and academic research worlds is a classic discussion. Academics have recently been suspected to address unmotivated problems because they do not manipulate the technologies that are at the core of their research activities. The importance of having an "industrial motivation" behind an academic research is reflected by a statistic: papers authored by at least one industrial researcher represent approximately half of accepted papers in the best conferences in operating systems (15 out of 32 for OSDI'11) and networking (16 out of 32 for Sigcomm'11). These papers monopolize the technical sessions related to new trends, especially datacenter and production network for OSDI, cloud computing and user measurement for Sigcomm.

    In these applied science areas, the best conferences accept papers addressing industry-relevant problems if and only if (i) authors demonstrate the timeliness and relevance of the problem, and (ii) authors carefully evaluate their proposed solutions.
    • problem motivation: a scientist who is only reading papers about a technology can hardly formulate a relevant important problem related to this technology. In order to have an accurate view of the problems faced by companies, a first idea is to spend time there as a visiting researcher, as it is promoted in Google. Another idea is to work with industrials in projects like  FP7 STREP project. I mean, actually work together, and not pretend working together.
    • solution evaluation: a NS2 simulation is no longer enough for a Sigcomm paper. Nowadays, some large-scale infrastructures give free access to scientists (for example Open Cirrus for a large data-center, Planet Lab for an Internet-scale network, Grid 5000 for a grid, Imagin'Lab for a 4G/LTE cellular network). There is no excuse to not test solutions over real infrastructures. However, the access to infrastructure is not sufficient, evaluations should also be based on realistic user patterns. Author of the excellent Hints and tips for Sigcomm authors claims "use realistic traffic models"! Besides using available real traces (for example the amazing network traces from Caida), the idea is again to leverage on a project collaboration with industrials that are able either to deploy a prototype on real clients, or to provide exclusive traces of their real clients.
    Hence, short-term focused collaborative projects are ideal if one wants to write well-motivated well-evaluated industry-relevant papers. But, in this case, why have I never been in position to submit a competitive paper to Sigcomm although I participated in many collaborative projects? Probably because:
    • some of my industrial partners were not really industrial. In large companies, R&D labs are frequently disconnected from the real operational teams, so researchers in these labs are unable to provide substantiated arguments about the criticality of the project, to successfully deploy a prototype, and even to obtain traces from their real clients.
    • in a consortium, every partner has its own agenda. Receiving fundings while minimizing efforts may be the only point all partners agree. I rarely feel that all partners share a strong commitment to make the project actually work. More frequently, the funding acceptance is considered as the final positive outcome, the project itself being only a pain.
    • the project work-plan does not include the writing of a scientific paper. Scientific production is usually seen as a dissemination activity, under the responsibility of an academic partner, although writing a top-class paper requires a precise planning of the contributions of every partner (including milestones and deliverables).
    Now that I understand why successful collaborative projects are critical and why my recent projects have (relatively) failed, I hope I will be able to leverage collaboration with industrials to do better research (a.k.a. write better papers).

    July 4, 2011

    My (disappointing) experience of attending a large conference

    Last month, I attended ICC at Kyoto. ICC is the kind of large-scale academic conference, where more than 1,400 people are expected to meet and collaborate for the sake of networking science. In the meantime, several other major conferences held in the so-called Federated Conferences at San Jose, which gathered approximately the same number of researchers. Several academic scientists have already reported their enthusiasm about such big events (or raised many positive thoughts).

    On my side, my experience was fairly negative. The technical and scientific discussions were rare, mostly because the conference scope is so large that the probability to chat with people sharing your scientific concerns is low. Actually, I have wondered what I was doing there during three quarters of conference time. Finally, I saw three reasons to attend big academic events:

    • awarding scientists: I think that scientists have been excellent students, that their commitment to excellence has again risen once they embrace an academic career, and that they are still not paid accordingly. Scientists do not receive bonuses in cash, however they frequently travel in wonderful places, with great banquet and rooms in palace. Conferencing is an award, which can be typically offered to a worthy PhD student. Similarly, professors do not hesitate to self-award with a full paid one-week conference (grants and funding allow traveling a lot, lets enjoy it). For those who like big hotels and international cities, big conferences are perfect.
    • meeting people from your local community: in a crowded amazingly large banquet, people first tend to cluster through languages or institutions. French chat with French, Chinese with Chinese, and members from University X with other members from University X. These "local cluster conversations" are easy to start (what plane did you take, how bad is the food in your hotel, how jetlagged are you, etc.). These local cluster conversations have at least one benefit: you have time to chat with people who you are used to meeting in local events without any chance to really discuss with. Therefore, a meeting in Japan is the place where you enhance your social network with people that live at less than 200 kilometers from your office, which is 10,000 kilometers away from Japan.
    • grenouilling during hallway conversations: it seems that the best translation of grenouiller is to plot.  In most multi-track conferences, many people prefer to stay in the lobby and do not attend talks. They are not wrong, because many scientific talks are actually bad, and I don't think that a series of talks is the best way to foster scientific conversations. However I am afraid that conversations in the lobby are not worth a trip of thousands of kilometers neither. A large part of conversations I heard or participated was related to research administration: what will be the next event-to-attend, am I in the Program Committee of next big event, what are the latest news about the next national funding call, where could my post-doc find a decent position, if I invite you in my steering committee, would you include me in your steering committee, what are the latest transfers in the academic world, etc.
    Probably because I expected some scientific enlightenments from meeting so many smart people, I have been disappointed. In particular, I definitely disagree with the scientists who argue for more maxi conferences. Next month, I will attend Sigcomm, which is a single-track reasonably-crowded (500 people) conference. Lets see if middle-size conferences are worth degrading my carbon footprint.

    April 28, 2011

    On the attractiveness of research institutions

    The job market for research tenured position is changing. First, as emphasized in a recent Nature issue, the number of graduating PhD is exploding. However the overall number of accepted papers in top-ranked conference is still desperately low. Hence, the vast majority of graduated PhD have a few minor publications and a h-index below 5. These young ambitious researchers are in the long tail of the scientists. Second, as illustrated by the closing of Intel Labs, the number of scientific jobs in private companies is dangerously decreasing. It seems that the future of research in private companies is about maintaining partnerships with universities and about outsourcing scientific studies to the right experts. In the meantime, the number of job openings in academic institutions is relatively stable. Third, academic institutions now hire scientists from all over the world. One of the consequences of the Shanghai ranking is that institutions do no longer focus on native candidates, and that many scientists look for jobs out of their native country.

    Please forget the most prestigious universities and the best young scientists. Both know how to match each other. Let's rather observe the second league: the thousands of more or less famous ambitious institutions, which want to attract the best scientists, and the thousands of more or less unknown ambitious scientists who want to join the best institutions. We are in a typical assignment problem, where institutions and scientists of similar ranking should agree.

    Many indicators have been proposed for the comparison of scientists. As well, many indicators or classification exist for ranking institutions for under-graduating students. But, I don't know how to measure the attractiveness of an academic institution for candidates to a tenure-track position.

    Of course, three criteria prevail:
    • the salary. The romantic vision of the scientist who does not care of money is wrong. Scientists are humans living in a capitalist world.
    • the location. It includes weather, probability that the family members can enjoy (employment, schools, etc.), cultural life, and so on.
    • the prestige of the institution. A scientist builds a career, her resume should maximize the number of famous entries and minimize unknown ones.
    Now, some more specific criteria include:
    • the number of free PhD students. Here, free means that the scientist does not have to produce any effort to have the guarantee that this number of PhD students will be under her advices in her lab.
    • the volume of teaching. It should not be too high because teaching must have no impact at all on the paper productivity. But it should not be too low because teaching is also a way to meet future PhD students.
    • the quality of students. Every scientist knows that bad students can be a significant waste of time, although great students can boost the productivity without much efforts.
    I am quite suspicious about the importance of having top-class colleagues within the same area. Ambitious scientists have their own research area, and a majority of them have their own agenda, without regards to other scientists. This thought leads me to another criteria:
    • the autonomy. From a scientific point of view, a researcher prefers to define her own research axis, and to write her own research proposals, without having to justify anything. From a more practical point of view, a researcher is likely to receive grants for her research, but this money goes first through her institution, which can constraint the expenses. Consequently, a scientist might be not free to buy her own equipment, not free to travel as she wishes, not free to set the salary of a post-doc, ... 
    Is there any other criteria? Of course, every researcher can introduce her own weight on this criteria, depending on her personal priority.

    It is now easy to analyze institutions. Typically, my current employer, Telecom Bretagne (member of Institut Telecom) :

    Salary
    7
    Good salary when you join, low increasing though
    Location
    4
    France is great, I do recommend Brest for a family with kids, but it is actually considered as a sub-attractive place in France
    Prestige
    3
    I am afraid that recent branding operations have significantly affected the reputation of the institution, nobody knows Institut Telecom
    Nb of free PhD students
    4
    No free PhD student at all, but it is not hard to obtain funding for one PhD student from the institution
    Volume teaching
    7
    It is highly negotiable with your colleagues, you can be involved in research-oriented project management rather than formal time-wasted courses
    Student quality
    8
    Very good engineering students, however very few are interested with research
    Autonomy
    8
    You are definitely free to do what you want, you are almost free to manage your own budget (personal, travel, equipment)
    Total
    41/70



    And Orange Labs, Issy-les-Moulineaux

    Salary
    8
    Good salary when you join, good opportunities to increase
    Location
    9
    Paris is one of the most attractive cities in the world
    Prestige
    5
    France Telecom was a strong actor of the research, Orange is a well-known brand, but the research center is no longer a key academic player
    Nb of free PhD students
    2
    No free PhD student at all, very hard to obtain it without significant efforts
    Volume teaching
    5
    No teaching at all, but not difficult to find some courses in nearby universities
    Student quality
    4
    No contact with students, but some students (not the best, though) might be interested with research internships
    Autonomy
    2
    Not free at all. You have to justify your research axis, and, worse, you have to justify any expense, even for funded projects
    Total
    35/70


    Now, we should create a website in order to gather ratings from several scientists, a kind of tripadvisor for scientific institutions.

    April 20, 2011

    inside the FP7 evaluation (last part): how to build funded proposals

    Don't expect any miracle from this post, and don't expect anything for posts with similar titles: there is no unique recipe for winning proposal. However, I can sketch a few tips that I will at least try to apply to my own proposals:
    • In the specific case of the FP7 STREP, every proposal is read by only three reviewers (no more no less) randomly picked among a set of very diverse reviewers, from experienced academics to young industrials (and vice versa) from various countries. Contrarily to the funding process of Google, there is no single reviewer target. Therefore, proposals should cover several "reading styles". Hopefully, there is no page limit, so don't hesitate to explain things several times in several ways.
    • The Criteria 1 ("Scientific and Technical Quality") can kill a proposal, but it can hardly make a winner. Your objective is to not give any opportunity for reviewer to criticize, so don't waste your energy there, just do a clean job. Recall that STREP is not about new scientific breakthrough, it is about incremental but sure progresses beyond State-of-the-Art. No reviewer can argue against a series of incremental loosely-consistent progresses in several domains. A bad note in Criteria 1 is more often due to a faulty or unconvincing workplan. Don't try to produce super-clever ambitious workplan, but describe things (scientific ideas and, more important, processes) that have 100% chances to be implemented. Revise your workplan a lot because some reviewers harshly track inconsistencies.
    • You have to differentiate, and the best place is Criteria 3 ("Exploitation and Dissemination"). The majority of proposals do not provide any market analysis, only claim standard dissemination plans, and describe very vague exploitation plans. The best proposals include real-world experimentations during the projects (which is actually appreciated), contain some partners that are actually involved in standardization processes, or have already identified some third-party companies to include in the proposal as external partners (or in a associated committee). The exploitation plan should be preferentially written at the beginning of the project, because it has a direct impact on the partners in the consortium, on the perimeter of the required scientific progresses and on the workplan (e.g. a real-world experimentation requires a specific work-package and a early prototype from technical work-packages). When the exploitation plan is strong, the whole project is fully consistent.
    I know that the usual process is definitely not this one. We usually work a lot on the scientific breakthroughs, trying to copy-paste endless bibliographic works written by students, to incorporate old scientific friends and to make the whole stuff as consistent as possible. Then, we let every partner write its own work package, and we argue a lot about the number of men/months and fundings. Finally, we have a few hours to write some crappy paragraphs about the exploitation. And we have wasted at least one month because the resulting proposal is all but a winner...

    March 14, 2011

    Inside the FP7 evaluation (part 2): the process

    The European Commission and committees of major academic conferences face a same challenge: select in a fair manner a dozen of proposals out of 100 (including 90 serious candidates). The process consists of three stages:
    1. reviewing the proposals. As I explained in my previous post, each proposal is a 100-pages long document, which is written by "artists". Three criteria are evaluated: one about the scientific soundness, one about the consortium quality and one about the exploitation. Every criteria is evaluated on a scale ranging from 0 to 5 where half-marks are accepted.
    2. reaching a consensus. For each proposal, five people (the three reviewers, one recorder and a moderator) meet during one hour. The goal is to reach a consensus, which should result in a unified text and a final score for every criteria. The role of the recorder is crucial. She has not read the proposal, but she looked at the reviews, so she knows the main trends. From the meeting discussion, she tries to extract some statements, then her text is revised "live" by the reviewers (and sometimes by the moderator). Wording is considered as important, so some sentences require up to 15 minutes to be accepted by every reviewer. In general, meetings are lively because some reviewers disagree, and it is common that reviewers actually argue. A consensus is reached in most meetings, but frequently in a unpleasant way because an enthusiastic reviewer has few chances to convince both other reviewers, and a positive-but-not-that-much consensus does not produce a winning project. In case of unreachable consensus, additional reviewers are invited to read the proposal. Eventually, a score is voted. 
    3. deciding. The panel committee meeting is like a program committee meeting except that a ranking is produced (even rejected proposals are ranked). The overall note ranges from 0 to 15, but of course, most proposals are between 8.5 and 13.5. There is a critical tie on a high score (around 13) because only a fraction of proposals having this score can be funded. A specific algorithm is used to break ties. In our case, proposals are ranked based on:
      1. the highest score in the Exploitation criteria, then, if tie again,
      2. the highest score in the Scientific criteria
      3. the largest ratio of industries
      4. the largest ratio of SMEs
      5. the largest ratio of partners from new member states of EU
      The role of the reviewers is actually marginal during this meeting: checking the consistency between final texts and final score for every proposal. Downgrading (or upgrading) a proposal after a quick cross-reading is very rare, and deserves a long agreement discussion from the panel. 
    The overall process suffers from a drawback: reviewers spend a lot of time on bad proposals. Every proposal, even the worst one, requires one hour of consensus meeting. Saving this time could let reviewers read a subset of the best proposals, and increase the quality of the final choice. In the panel meeting, a long time is also wasted on revising the text for every proposal, even the ones that will not be funded, although panelists do not have sufficient time to discuss the borderline proposals.

    The consensus part is funny. The overall result of the consensus meetings is rarely a "blind union" of three independent reviews (as it is done in most conferences, like averaging three scores). For example, three reviewers adopted a 4.0 for the scientific part (which is a very good score) but they reached a consensus with a score of 2.5 (which is below the threshold) because the flaws they identified were complementary, or because they discovered that they share an overall lack of excitement about the proposal, so they took time to detect actual flaws justifying the reject.

    The overall process is fair, and there is no way to express any subjective opinion, like this topic is funnier than the other, or these guys should be assisted because their country bankrupts, or I don't like this crappy acronym

    March 1, 2011

    inside the FP7 evaluation (part 1?)

    I am involved in the process of project evaluation for the European Commission for the first time. Some selected remarks:
    • there is an art of writing proposals. Scientists know that the art of writing academic papers has become a key skill in the modern science battlefield. The art of writing proposal is also widely admitted, but I had never faced it. Now I know. The top proposals I reviewed have a lot in common, including the approach and the style. Generally, these projects are about an exciting but obscured concept (vaporware?), supported by a very standard research from high-h-ranked scientists (business as usual), in cooperation with fresh SMEs (old students), and the classic large company (whose role is... hmm, well, to be there). These proposals have probably been powerpoint in a previous life: objectives are presented in a bullet-mode emphatic way, the proposal is un-verbose (so less risk of inconsistency), every page contains a figure or a table. As can be expected, the workpackage organization is perfect with an ideal balance of man-month by workpackage and by partners. It is difficult to know the future of such project. The academic work is probably already under submission. Several web revolutions will occur until the end of the project. The final software will probably fail, because it will be coded by un-managed students in University. However, in the evaluation form of European Commission, such a proposal deserves a "check" for every critical parameter, so at the end, they have good chances to win.
    • it is innovation, it is not about research. For those who had doubts. The best proposals are ambitious. Most of them include some attractive real-world experimentations, which requires committing a lot of developers. As the overall funding is constrained, the research-oriented demand is minimal. Moreover, as previously said, no inconsistency is tolerable, therefore every dozen of claimed man-month should be justified, and related with the remaining of the project, which is very short-term. Therefore, the scientific topic of every participating "researcher" is approximately defined in advance. Is it research? Of course not, it is the so-called innovation by research. I tend to be in favor of such early development project, however the other funding agencies (local area and national) follow the same objective, so is there a way to make un-purposed research? And why the hell are there so few innovations by research from Europe although it is already the seventh similar program?
    • the reviewing process is very short, but well paid. I detected one bad consequence from being a paid reviewer: reviewers have incentives to evaluate more projects, although they have no time to evaluate them carefully. Delays are very tight: I had seven projects to review in less than twelve days, each project being a hundred pages document, which details a three years long study by a consortium ranging from six to twelve partners. Moreover, the quality of the proposals was excellent, even for the worst one. Hence, one half day for a complete reviewing is minimum for a (slow?) young scientist like me. The risk is to produce a quick evaluation based on the strategy of "killing a project for any small detail". My personal reviewing strategy (I saw several people doing the same) was to first have a very quick first pass on the whole document, then to go into details once the overall concept was clear. In this context, no inconsistency is tolerable, it is better to have only one proposal writer, preferentially someone who... now come back to the first point of this post.
    I will probably have more to say after my week at Brussels for the final decision.